Files
PFVCluster/STATUS.md
T
mrcharles cd1c2f7332 docs(status): record operator schedule + remote vs onsite work split
Operator onsite Monday 0700. Until then: software/docs/OAM only (Sunday
off). Switch changes + physical cable work deferred to Monday onsite.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 23:03:48 -05:00

12 KiB
Raw Blame History

Project Status

Human read-only. Agents maintain this file automatically after each work session. Do not edit by hand — the next agent run will overwrite it.

Last updated: 2026-07-30 by Crush (GLM-5.2) — full ground-truth audit (network+storage+compute)

Current authoritative audit: proxmox/docs/AUDIT-2026-07-30.md — network topology, switch LAG configs, corrected storage topology, VM config audit, fleet consistency, cross-rack NFS root-cause analysis. Capacity targets are codified in proxmox/docs/CAPACITY-GOALS.md: 80% RAM steady-state, ~50% idle CPU.

Current State: STABLE

The merged PFVCluster repo is fully operational across provisioning, DNS infrastructure, Proxmox cluster ops, and k8s control plane.

Completed Work

Server Provisioning (validated on sectestbed-sandbox)

  • SetupNewSystem.sh deploys end-to-end to rc=0 (Debian 13 trixie)
  • Security hardening: SSH, SCAP-STIG, Wazuh, 2FA (SSH+Cockpit+Webmin)
  • NTP: redundant pfv-netinfra-01/02 (192.168.3.252/253), synced stratum 3
  • DNS resolv.conf: managed static file pointing at netinfra pair
  • Test suite: 5 tests (framework, safe-download, 2fa, https, system-req)
  • VM validation harness: git-based deploy + auto-rollback + guest-agent access

DNS Cluster (pfv-netinfra-01/02)

  • Production Technitium config replicated from tailscale-router (read-only)
  • 124 zones on both nodes (knel.net + reverse DNS)
  • pfv-netinfra-01 = PRIMARY, pfv-netinfra-02 = SECONDARY
  • Zone replication via rsync systemd timer (every 60s)
  • Credentials + 2FA replicated identically to production
  • Both LAN IPs resolve knel.net device names + recurse externally
  • DHCP migrated from pfv-netboot → netinfra-01/02 (ISC failover-peer, primary/secondary). Both normal, 333 active leases load-balanced. DNS+NTP now handed out as 252/253. Netboot DHCP stopped + disabled. Webmin DHCP module on both nodes.

Proxmox Cluster Ops

  • 5 of 7 hosts fully performance-tuned (tsys1/3/6/7/9)
  • Full ground-truth audit 2026-07-30 (network + storage + compute + switches)
  • Cross-rack NFS root cause identified: core-sw01 LAG hash=layer-2-3 (no L4), g16/g17 cable mismatch (3 of 4 links active), static mode (no LACP)
  • tsys4/tsys5 storage changes documented: new NVMe (CT500P1 500GB), relocated SK hynix SSD, D3 changed to USB HDD, /mnt/albert + /mnt/backup added
  • Console symlink bug fixed: udev rule singular→plural mismatch broke all 7 console ports after reboot
  • Portable audit tooling built: probe-storage.sh, probe-network.sh, conman-console.py, snmp-switch-audit.py (AGPLv3-friendly, config-driven)
  • Memory goal (≤80% actual): PASS fleet-wide — highest is tsys9 at 69%
  • CPU ~50%-idle goal: fleet idles 0-5% (under-utilized; chase via workloads, not hardware)
  • tsys4: PVE 9.1.5 (behind fleet 9.2.5), needs upgrade
  • tsys5: NVMe + SK hynix SSD installed but NOT configured
  • tsys2: pending rebuild from Win10 to Proxmox

Kubernetes Control Plane (k3s HA — LIVE)

  • 3-node k3s HA control plane deployed: cnode1/2/3 (v1.36.2+k3s1, embedded etcd)
  • All traffic over Tailscale IPs — no LAN IPs in node status or certs
  • All 3 cnodes tainted control-plane:NoSchedule (zero user workloads)
  • 13/13 health checks pass (verify.sh): nodes Ready, etcd quorum, Tailscale IPs, CoreDNS, API server, workload isolation
  • Scripts in k8s/: wipe, install-cp, join-servers, post-setup, verify
  • Kubeconfig saved to ~/.kube/config.pfv-k8s (gitignored, embedded certs)
  • Workers (wnodes) not yet joined to this cluster
  • Distro decision: k3s chosen for regular R&D cluster. Talos docs preserved in k8s/docs/ for future ITAR/classified cluster.

PDU Management (powerman on pfv-tsys1 — LIVE)

  • Cyclades AlterPath PM10i (10 outlets) managed via powerman over serial
  • USB-DB9 adapter (Prolific pl2303) with stable udev symlink /dev/cyclades-pm10
  • powermand listening on 127.0.0.1:10101 + Tailscale 100.121.189.98:10101
  • All 10 outlets defined as outlet-1 through outlet-10
  • Validated: outlet 10 cycled off → on (8/8 test checks passed)
  • Rename outlets to match physical devices (Friday onsite)

UPS Management (NUT on pfv-tsys1 — LIVE)

  • APC Smart-UPS C 1500 fully reporting via NUT usbhid-ups + APC HID 0.100
  • upsd listening on Tailscale 100.121.189.98:3493 (port 3493)
  • Local upsmon (master) configured for graceful hypervisor shutdown
  • Home Assistant credentials generated (homeassistant user in upsd.users)
  • Scripts in ups/: discover, setup (idempotent), status
  • Home Assistant NUT integration configured via REST API (ups/setup-ha-nut.sh)
  • Key fix: upsd now listens on LAN IP (192.168.3.11) in addition to Tailscale — HAOS runs Tailscale as an isolated add-on, so the HA container cannot route to Tailscale IPs. Shared vmbr0 bridge provides instant LAN path.
  • Key fix: udev rule covers BOTH usb and hidraw subsystems (driver opens raw /dev/bus/usb after setuid drop, not /dev/hidraw)
  • Tripp Lite UPS blocked: driver matches TrippLite HID 0.85 subdriver but fails reading 878-byte HID Report Descriptor (EAGAIN via USB hub). Needs physical reseat to direct motherboard USB port or cable replacement. Driver masked to prevent restart-loop spam.
  • Home Assistant NUT integration: configured and live (3 sensors: battery charge, status, status data). This APC (FW 02.2, mfg 2012) does not report ups.load over USB HID — confirmed via both NUT and apcupsd tests. No wattage for Energy Dashboard; smart plug on UPS output is the practical path. apcupsd package conflicts with nut-server (mutually exclusive).

Console Management (ser2net telnet(rfc2217) + conman on pfv-tsys4 — LIVE)

  • 7 network switch/router consoles managed via ser2net + conman
  • USB enumeration problem SOLVED: udev rules pin each adapter by ID_PATH (physical USB port topology) to stable /dev/consoles/<name> symlinks that survive reboot regardless of enumeration order
  • ser2net exposes all 7 consoles on telnet(rfc2217) TCP ports (2001-2007) bound to Tailscale IP only (100.70.77.93)
  • conman connects to TCP ports for logging + multiplexing (7 log files active in /var/log/conman/)
  • Both ser2net + conmand enabled via systemd (survive reboot)
  • conmand systemd unit created (Debian package doesn't ship one)
  • conmand remote access enabled (loopback=off, port 7890)
  • Old /root/conmap + manual screen workflow replaced
  • Stair-stepping fix: ser2net accepter changed from raw tcp to telnet(rfc2217) so conman's telnet NVT negotiates binary mode properly
  • Validated 6/7 consoles working via conman from workstation (core-sw01, tor3-mgmt, tor3-stor, rrinfra-rtr, r2-tor-top, subodev-torsw)
  • pfv-r2-sw (port 2007): DEAD — zero bytes at all baud rates (1200-115200), adapter present and mapped correctly (/dev/ttyUSB3). Physical issue: device powered off, console cable disconnected, or console port disabled in device config. Requires onsite investigation.

Ansible AWX (tsys-awx — LIVE)

  • AWX 24.6.1 deployed via AWX Operator 2.19.1 on k3s single-node
  • Accessible at http://tsys-awx.knel.net (port 80, LoadBalancer via k3s ServiceLB)
  • Admin login: admin / REDACTED_PASSWORD
  • PostgreSQL 15 (bundled, 8 GiB PVC on local-path)
  • VM resized: 32→60 GB disk, swap converted to swapfile
  • Scripts in awx/: install-k3s, deploy-awx, verify-awx
  • HTTPS/TLS not yet configured (currently HTTP on port 80)
  • Execution environments not yet customized

Repo Merge

  • KNELServerBuild merged into PFVCluster (history preserved)
  • Directory structure reorganized (provisioning/, tests/, perf/, docs/)
  • All docs gardened: links fixed, stale refs removed, tailscale.md updated

Known Issues

Issue Impact Status
core-sw01 LAG hash = layer-2-3 NFS cross-rack capped at 1 Gbps per host pair Fix: change to layer-2-3-4 (switch change, pending approval)
core-sw01 ch1 g16/g17 mismatch Only 3 of 4 cross-rack links active Fix: change member list g(13-15,17)→g(13-16)
Cross-rack trunk uses mode on No LACP failure detection Fix: change to mode active (both switches)
tsys5 bond0 degraded (1 of 2 slaves) Single link, no redundancy Physical: 2nd ethernet cable needed
tsys5 NVMe + SSD not configured Best local storage unused Config: mount + add to Proxmox storage.cfg
D2 carries 13 running VMs IOPS hotspot persists Migrate to S2/S3/new disks
tsys4 PVE 9.1.5 Behind fleet (9.2.5) Upgrade needed
tsys5 SSD 93% full Samsung 860 PRO nearly exhausted Offload to new SK hynix SSD
Port errors on core-sw01 g9:2902, g38:590, g26:73 inErr Investigate CRC/cable
tsys5 is single-CPU CPU2 socket unpopulated Consider seating spare LGA771 Xeon
Fleet CPU 0-5% idle Misses ~50%-idle utilization goal Consolidate baseline; grow via workloads

Pending (next session priorities)

Full audit with findings + recommendations is in proxmox/docs/AUDIT-2026-07-30.md.

Awaiting operator review of AUDIT-2026-07-30 before executing any fixes.

Switch changes (need approval):

  • Fix core-sw01 ch1 member: g(13-15,17) → g(13-16) [restores 4th link]
  • Fix core-sw01 LAG hash: layer-2-3 → layer-2-3-4 [unlocks multi-link NFS]
  • Convert trunk to LACP: mode on → mode active [failure detection]
  • Remove stale tsys8 config from tor3-stor (ch4/g13/g14)

Host-side changes (no switch access needed):

  • Mount + configure tsys5 NVMe (CT500P1 500GB) as LVM-thin
  • Mount + configure tsys5 SK hynix SSD (512GB, relocated from tsys4)
  • Clean stale S3/S4 exports + storage.cfg on tsys5
  • Upgrade tsys4 to PVE 9.2.5 / kernel 7.0.14
  • Apply apply-tunings.sh to tsys4/5 (nconnect, bond hash, rmem/wmem)
  • Audit VM cache modes (writethrough → writeback for NFS)

Other:

  • Schedule: Operator onsite Monday 0700. Until then: software fixes, documentation, and OAM only (no physical access). Sunday off. Saturday = OAM Day (2026-08-01).
  • Remote-capable (nowSaturday): tsys5 NVMe/SSD mount+config, fleet tuning (tsys4/5 nconnect/bond hash/rmem), VM config audit, OAM Day prep, documentation
  • Onsite-only (Monday+): switch config changes (core-sw01 ch1 hash+member, tor3-stor LACP conversion), tsys5 2nd ethernet cable, physical hardware work
  • Saturday OAM Day (2026-08-01): 100% instrumentation coverage — Beszel (metrics) on all hosts, Uptime Kuma (up/down) → Pushover alerts, LibreNMS (SNMP) for switches, Home Assistant (power/temp) for physical infra, centralized syslog (tool TBD)
  • Investigate core-sw01 port errors (g9/g38/g26)
  • k8s: join wnodes, label by workload class
  • k8s deferred: ETL tooling, HPC scheduler, vcluster policy

Infrastructure Summary

Component Details
Proxmox hosts 7 standalone PVE, managed via PDM (tsys8 retired)
Storage servers tsys4 (NFS: D1/D2/D3/D5 + /mnt/albert + /mnt/backup), tsys5 (NFS: S1/S2 + local SSD + new NVMe unconfigured)
Switches core-sw01 (PowerConnect 5448, rack5), tor3-stor/mgmt (PowerConnect 5324, rack3), + 1 unidentified (192.168.0.8)
DNS primary pfv-netinfra-01 (192.168.3.252) — Technitium + Pi-hole
DNS secondary pfv-netinfra-02 (192.168.3.253) — Technitium + Pi-hole
DNS production tailscale-router (read-only source of truth)
NTP pfv-netinfra-01/02 (redundant, LAN IPs, stratum 2/3)
Sandbox VM sectestbed-sandbox (VMID 6000 on pfv-tsys5)
Backup Proxmox Backup Server (PBS)
k8s control plane 3-node k3s HA (cnode1/2/3), all traffic over Tailscale
PDU Cyclades PM10i via powerman on pfv-tsys1 (port 10101)
AWX tsys-awx: AWX 24.6.1 on k3s (http://tsys-awx.knel.net:80)
Console 7 switch consoles via ser2net telnet(rfc2217)+conman on pfv-tsys4 (conmand:7890 on Tailscale)
UPS APC Smart-UPS C 1500 via NUT on pfv-tsys1 (upsd:3493 on Tailscale)
Production Cloudron VPS, Reston VA (this cluster is R&D only)