Operator onsite Monday 0700. Until then: software/docs/OAM only (Sunday
off). Switch changes + physical cable work deferred to Monday onsite.
💘 Generated with Crush
Assisted-by: Crush:glm-5.2
12 KiB
Project Status
Human read-only. Agents maintain this file automatically after each work session. Do not edit by hand — the next agent run will overwrite it.
Last updated: 2026-07-30 by Crush (GLM-5.2) — full ground-truth audit (network+storage+compute)
Current authoritative audit:
proxmox/docs/AUDIT-2026-07-30.md— network topology, switch LAG configs, corrected storage topology, VM config audit, fleet consistency, cross-rack NFS root-cause analysis. Capacity targets are codified inproxmox/docs/CAPACITY-GOALS.md: 80% RAM steady-state, ~50% idle CPU.
Current State: STABLE
The merged PFVCluster repo is fully operational across provisioning, DNS infrastructure, Proxmox cluster ops, and k8s control plane.
Completed Work
Server Provisioning (validated on sectestbed-sandbox)
- SetupNewSystem.sh deploys end-to-end to rc=0 (Debian 13 trixie)
- Security hardening: SSH, SCAP-STIG, Wazuh, 2FA (SSH+Cockpit+Webmin)
- NTP: redundant pfv-netinfra-01/02 (192.168.3.252/253), synced stratum 3
- DNS resolv.conf: managed static file pointing at netinfra pair
- Test suite: 5 tests (framework, safe-download, 2fa, https, system-req)
- VM validation harness: git-based deploy + auto-rollback + guest-agent access
DNS Cluster (pfv-netinfra-01/02)
- Production Technitium config replicated from tailscale-router (read-only)
- 124 zones on both nodes (knel.net + reverse DNS)
- pfv-netinfra-01 = PRIMARY, pfv-netinfra-02 = SECONDARY
- Zone replication via rsync systemd timer (every 60s)
- Credentials + 2FA replicated identically to production
- Both LAN IPs resolve knel.net device names + recurse externally
- DHCP migrated from pfv-netboot → netinfra-01/02 (ISC failover-peer, primary/secondary). Both
normal, 333 active leases load-balanced. DNS+NTP now handed out as 252/253. Netboot DHCP stopped + disabled. Webmin DHCP module on both nodes.
Proxmox Cluster Ops
- 5 of 7 hosts fully performance-tuned (tsys1/3/6/7/9)
- Full ground-truth audit 2026-07-30 (network + storage + compute + switches)
- Cross-rack NFS root cause identified: core-sw01 LAG hash=layer-2-3 (no L4), g16/g17 cable mismatch (3 of 4 links active), static mode (no LACP)
- tsys4/tsys5 storage changes documented: new NVMe (CT500P1 500GB), relocated SK hynix SSD, D3 changed to USB HDD, /mnt/albert + /mnt/backup added
- Console symlink bug fixed: udev rule singular→plural mismatch broke all 7 console ports after reboot
- Portable audit tooling built: probe-storage.sh, probe-network.sh, conman-console.py, snmp-switch-audit.py (AGPLv3-friendly, config-driven)
- Memory goal (≤80% actual): PASS fleet-wide — highest is tsys9 at 69%
- CPU ~50%-idle goal: fleet idles 0-5% (under-utilized; chase via workloads, not hardware)
- tsys4: PVE 9.1.5 (behind fleet 9.2.5), needs upgrade
- tsys5: NVMe + SK hynix SSD installed but NOT configured
- tsys2: pending rebuild from Win10 to Proxmox
Kubernetes Control Plane (k3s HA — LIVE)
- 3-node k3s HA control plane deployed: cnode1/2/3 (v1.36.2+k3s1, embedded etcd)
- All traffic over Tailscale IPs — no LAN IPs in node status or certs
- All 3 cnodes tainted
control-plane:NoSchedule(zero user workloads) - 13/13 health checks pass (verify.sh): nodes Ready, etcd quorum, Tailscale IPs, CoreDNS, API server, workload isolation
- Scripts in
k8s/: wipe, install-cp, join-servers, post-setup, verify - Kubeconfig saved to
~/.kube/config.pfv-k8s(gitignored, embedded certs) - Workers (wnodes) not yet joined to this cluster
- Distro decision: k3s chosen for regular R&D cluster. Talos docs
preserved in
k8s/docs/for future ITAR/classified cluster.
PDU Management (powerman on pfv-tsys1 — LIVE)
- Cyclades AlterPath PM10i (10 outlets) managed via powerman over serial
- USB-DB9 adapter (Prolific pl2303) with stable udev symlink
/dev/cyclades-pm10 - powermand listening on
127.0.0.1:10101+ Tailscale100.121.189.98:10101 - All 10 outlets defined as
outlet-1throughoutlet-10 - Validated: outlet 10 cycled off → on (8/8 test checks passed)
- Rename outlets to match physical devices (Friday onsite)
UPS Management (NUT on pfv-tsys1 — LIVE)
- APC Smart-UPS C 1500 fully reporting via NUT usbhid-ups + APC HID 0.100
- upsd listening on Tailscale
100.121.189.98:3493(port 3493) - Local upsmon (master) configured for graceful hypervisor shutdown
- Home Assistant credentials generated (
homeassistantuser in upsd.users) - Scripts in
ups/: discover, setup (idempotent), status - Home Assistant NUT integration configured via REST API (
ups/setup-ha-nut.sh) - Key fix: upsd now listens on LAN IP (192.168.3.11) in addition to Tailscale — HAOS runs Tailscale as an isolated add-on, so the HA container cannot route to Tailscale IPs. Shared vmbr0 bridge provides instant LAN path.
- Key fix: udev rule covers BOTH
usbandhidrawsubsystems (driver opens raw/dev/bus/usbafter setuid drop, not/dev/hidraw) - Tripp Lite UPS blocked: driver matches TrippLite HID 0.85 subdriver but fails reading 878-byte HID Report Descriptor (EAGAIN via USB hub). Needs physical reseat to direct motherboard USB port or cable replacement. Driver masked to prevent restart-loop spam.
- Home Assistant NUT integration: configured and live (3 sensors:
battery charge, status, status data). This APC (FW 02.2, mfg 2012) does not
report
ups.loadover USB HID — confirmed via both NUT and apcupsd tests. No wattage for Energy Dashboard; smart plug on UPS output is the practical path. apcupsd package conflicts with nut-server (mutually exclusive).
Console Management (ser2net telnet(rfc2217) + conman on pfv-tsys4 — LIVE)
- 7 network switch/router consoles managed via ser2net + conman
- USB enumeration problem SOLVED: udev rules pin each adapter by
ID_PATH (physical USB port topology) to stable
/dev/consoles/<name>symlinks that survive reboot regardless of enumeration order - ser2net exposes all 7 consoles on telnet(rfc2217) TCP ports (2001-2007)
bound to Tailscale IP only (
100.70.77.93) - conman connects to TCP ports for logging + multiplexing (7 log
files active in
/var/log/conman/) - Both ser2net + conmand enabled via systemd (survive reboot)
- conmand systemd unit created (Debian package doesn't ship one)
- conmand remote access enabled (loopback=off, port 7890)
- Old
/root/conmap+ manualscreenworkflow replaced - Stair-stepping fix: ser2net accepter changed from raw
tcptotelnet(rfc2217)so conman's telnet NVT negotiates binary mode properly - Validated 6/7 consoles working via conman from workstation (core-sw01, tor3-mgmt, tor3-stor, rrinfra-rtr, r2-tor-top, subodev-torsw)
- pfv-r2-sw (port 2007): DEAD — zero bytes at all baud rates (1200-115200), adapter present and mapped correctly (/dev/ttyUSB3). Physical issue: device powered off, console cable disconnected, or console port disabled in device config. Requires onsite investigation.
Ansible AWX (tsys-awx — LIVE)
- AWX 24.6.1 deployed via AWX Operator 2.19.1 on k3s single-node
- Accessible at
http://tsys-awx.knel.net(port 80, LoadBalancer via k3s ServiceLB) - Admin login:
admin/REDACTED_PASSWORD - PostgreSQL 15 (bundled, 8 GiB PVC on local-path)
- VM resized: 32→60 GB disk, swap converted to swapfile
- Scripts in
awx/: install-k3s, deploy-awx, verify-awx - HTTPS/TLS not yet configured (currently HTTP on port 80)
- Execution environments not yet customized
Repo Merge
- KNELServerBuild merged into PFVCluster (history preserved)
- Directory structure reorganized (provisioning/, tests/, perf/, docs/)
- All docs gardened: links fixed, stale refs removed, tailscale.md updated
Known Issues
| Issue | Impact | Status |
|---|---|---|
| core-sw01 LAG hash = layer-2-3 | NFS cross-rack capped at 1 Gbps per host pair | Fix: change to layer-2-3-4 (switch change, pending approval) |
| core-sw01 ch1 g16/g17 mismatch | Only 3 of 4 cross-rack links active | Fix: change member list g(13-15,17)→g(13-16) |
Cross-rack trunk uses mode on |
No LACP failure detection | Fix: change to mode active (both switches) |
| tsys5 bond0 degraded (1 of 2 slaves) | Single link, no redundancy | Physical: 2nd ethernet cable needed |
| tsys5 NVMe + SSD not configured | Best local storage unused | Config: mount + add to Proxmox storage.cfg |
| D2 carries 13 running VMs | IOPS hotspot persists | Migrate to S2/S3/new disks |
| tsys4 PVE 9.1.5 | Behind fleet (9.2.5) | Upgrade needed |
| tsys5 SSD 93% full | Samsung 860 PRO nearly exhausted | Offload to new SK hynix SSD |
| Port errors on core-sw01 | g9:2902, g38:590, g26:73 inErr | Investigate CRC/cable |
| tsys5 is single-CPU | CPU2 socket unpopulated | Consider seating spare LGA771 Xeon |
| Fleet CPU 0-5% idle | Misses ~50%-idle utilization goal | Consolidate baseline; grow via workloads |
Pending (next session priorities)
Full audit with findings + recommendations is in
proxmox/docs/AUDIT-2026-07-30.md.
Awaiting operator review of AUDIT-2026-07-30 before executing any fixes.
Switch changes (need approval):
- Fix core-sw01 ch1 member: g(13-15,17) → g(13-16) [restores 4th link]
- Fix core-sw01 LAG hash: layer-2-3 → layer-2-3-4 [unlocks multi-link NFS]
- Convert trunk to LACP: mode on → mode active [failure detection]
- Remove stale tsys8 config from tor3-stor (ch4/g13/g14)
Host-side changes (no switch access needed):
- Mount + configure tsys5 NVMe (CT500P1 500GB) as LVM-thin
- Mount + configure tsys5 SK hynix SSD (512GB, relocated from tsys4)
- Clean stale S3/S4 exports + storage.cfg on tsys5
- Upgrade tsys4 to PVE 9.2.5 / kernel 7.0.14
- Apply apply-tunings.sh to tsys4/5 (nconnect, bond hash, rmem/wmem)
- Audit VM cache modes (writethrough → writeback for NFS)
Other:
- Schedule: Operator onsite Monday 0700. Until then: software fixes, documentation, and OAM only (no physical access). Sunday off. Saturday = OAM Day (2026-08-01).
- Remote-capable (now–Saturday): tsys5 NVMe/SSD mount+config, fleet tuning (tsys4/5 nconnect/bond hash/rmem), VM config audit, OAM Day prep, documentation
- Onsite-only (Monday+): switch config changes (core-sw01 ch1 hash+member, tor3-stor LACP conversion), tsys5 2nd ethernet cable, physical hardware work
- Saturday OAM Day (2026-08-01): 100% instrumentation coverage — Beszel (metrics) on all hosts, Uptime Kuma (up/down) → Pushover alerts, LibreNMS (SNMP) for switches, Home Assistant (power/temp) for physical infra, centralized syslog (tool TBD)
- Investigate core-sw01 port errors (g9/g38/g26)
- k8s: join wnodes, label by workload class
- k8s deferred: ETL tooling, HPC scheduler, vcluster policy
Infrastructure Summary
| Component | Details |
|---|---|
| Proxmox hosts | 7 standalone PVE, managed via PDM (tsys8 retired) |
| Storage servers | tsys4 (NFS: D1/D2/D3/D5 + /mnt/albert + /mnt/backup), tsys5 (NFS: S1/S2 + local SSD + new NVMe unconfigured) |
| Switches | core-sw01 (PowerConnect 5448, rack5), tor3-stor/mgmt (PowerConnect 5324, rack3), + 1 unidentified (192.168.0.8) |
| DNS primary | pfv-netinfra-01 (192.168.3.252) — Technitium + Pi-hole |
| DNS secondary | pfv-netinfra-02 (192.168.3.253) — Technitium + Pi-hole |
| DNS production | tailscale-router (read-only source of truth) |
| NTP | pfv-netinfra-01/02 (redundant, LAN IPs, stratum 2/3) |
| Sandbox VM | sectestbed-sandbox (VMID 6000 on pfv-tsys5) |
| Backup | Proxmox Backup Server (PBS) |
| k8s control plane | 3-node k3s HA (cnode1/2/3), all traffic over Tailscale |
| PDU | Cyclades PM10i via powerman on pfv-tsys1 (port 10101) |
| AWX | tsys-awx: AWX 24.6.1 on k3s (http://tsys-awx.knel.net:80) |
| Console | 7 switch consoles via ser2net telnet(rfc2217)+conman on pfv-tsys4 (conmand:7890 on Tailscale) |
| UPS | APC Smart-UPS C 1500 via NUT on pfv-tsys1 (upsd:3493 on Tailscale) |
| Production | Cloudron VPS, Reston VA (this cluster is R&D only) |