Files
PFVCluster/STATUS.md
T
mrcharles fa0c77fb58 fix(console): switch conman from ser2net-telnet to direct serial access
The ser2net layer between conman and the serial devices was causing
stair-stepping in terminal output. conman's dev="host:port" uses telnet
protocol, but ser2net's accepter was raw TCP (no telnet mode). The telnet
NVT state machine in conman was stripping bare CR characters from device
output — particularly from Dell switches that use old-style \n\r (LF+CR)
line endings instead of standard \r\n. In telnet, a bare \r not followed
by \n or NUL is non-compliant and gets dropped, leaving bare \n that
causes stair-stepping in the raw-mode terminal.

Fix: conman now opens serial devices directly via the stable udev
symlinks (/dev/consoles/<name>) with seropts, eliminating the telnet
layer entirely. ser2net is stopped and disabled but remains installed
for emergency TCP access (documented workflow: stop conmand, start
ser2net, use telnet, then reverse).

Architecture change:
  Before: device → serial → ser2net (raw TCP) → conman (telnet NVT) → terminal
  After:  device → serial → conman (direct) → terminal

Verified: 7/7 serial devices held by conmand, 7/7 log files capturing,
all \r bytes preserved in terminal output (confirmed via PTY capture).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 18:29:31 -05:00

8.3 KiB

Project Status

Human read-only. Agents maintain this file automatically after each work session. Do not edit by hand — the next agent run will overwrite it.

Last updated: 2026-07-29 by Crush (GLM-5.2)

Capacity targets are now codified in proxmox/docs/CAPACITY-GOALS.md: 80% RAM steady-state, ~50% idle CPU. The current placement analysis lives in proxmox/docs/AUDIT-2026-07-29.md.

Current State: STABLE

The merged PFVCluster repo is fully operational across provisioning, DNS infrastructure, Proxmox cluster ops, and k8s control plane.

Completed Work

Server Provisioning (validated on sectestbed-sandbox)

  • SetupNewSystem.sh deploys end-to-end to rc=0 (Debian 13 trixie)
  • Security hardening: SSH, SCAP-STIG, Wazuh, 2FA (SSH+Cockpit+Webmin)
  • NTP: redundant pfv-netinfra-01/02 (192.168.3.252/253), synced stratum 3
  • DNS resolv.conf: managed static file pointing at netinfra pair
  • Test suite: 5 tests (framework, safe-download, 2fa, https, system-req)
  • VM validation harness: git-based deploy + auto-rollback + guest-agent access

DNS Cluster (pfv-netinfra-01/02)

  • Production Technitium config replicated from tailscale-router (read-only)
  • 124 zones on both nodes (knel.net + reverse DNS)
  • pfv-netinfra-01 = PRIMARY, pfv-netinfra-02 = SECONDARY
  • Zone replication via rsync systemd timer (every 60s)
  • Credentials + 2FA replicated identically to production
  • Both LAN IPs resolve knel.net device names + recurse externally
  • DHCP migrated from pfv-netboot → netinfra-01/02 (ISC failover-peer, primary/secondary). Both normal, 333 active leases load-balanced. DNS+NTP now handed out as 252/253. Netboot DHCP stopped + disabled. Webmin DHCP module on both nodes.

Proxmox Cluster Ops

  • 5 of 7 hosts fully performance-tuned (tsys1/3/6/7/9)
  • Fleet audit refreshed 2026-07-29 (consolidated, storage-corrected)
  • Memory goal (≤80% actual) met fleet-wide — highest host is tsys1 at 73%
  • Capacity goals codified (80% RAM / ~50% idle CPU + workload model)
  • Storage topology corrected: local-nonprod = S1 disk (not separate); per-host storage.cfg
  • cnode5 (706) deleted; cnode4 (601) stopped (delete pending)
  • tsys6 S1 typo + tsys9 missing S1 — both FIXED by operator (verified: 8/8 NFS stores active on both)
  • CPU ~50%-idle goal: fleet idles 0-5% (under-utilized; chase via workloads, not hardware)
  • tsys4: blocked on PCIe NIC + RAM install
  • tsys5: blocked on 2nd ethernet cable + NVMe install; also single-CPU (CPU2 socket empty)
  • tsys2: pending rebuild from Win10 to Proxmox

Kubernetes Control Plane (k3s HA — LIVE)

  • 3-node k3s HA control plane deployed: cnode1/2/3 (v1.36.2+k3s1, embedded etcd)
  • All traffic over Tailscale IPs — no LAN IPs in node status or certs
  • All 3 cnodes tainted control-plane:NoSchedule (zero user workloads)
  • 13/13 health checks pass (verify.sh): nodes Ready, etcd quorum, Tailscale IPs, CoreDNS, API server, workload isolation
  • Scripts in k8s/: wipe, install-cp, join-servers, post-setup, verify
  • Kubeconfig saved to ~/.kube/config.pfv-k8s (gitignored, embedded certs)
  • Workers (wnodes) not yet joined to this cluster
  • Distro decision: k3s chosen for regular R&D cluster. Talos docs preserved in k8s/docs/ for future ITAR/classified cluster.

PDU Management (powerman on pfv-tsys1 — LIVE)

  • Cyclades AlterPath PM10i (10 outlets) managed via powerman over serial
  • USB-DB9 adapter (Prolific pl2303) with stable udev symlink /dev/cyclades-pm10
  • powermand listening on 127.0.0.1:10101 + Tailscale 100.121.189.98:10101
  • All 10 outlets defined as outlet-1 through outlet-10
  • Validated: outlet 10 cycled off → on (8/8 test checks passed)
  • Rename outlets to match physical devices (Friday onsite)

Console Management (conman direct-serial on pfv-tsys4 — LIVE)

  • 7 network switch/router consoles managed via conman
  • USB enumeration problem SOLVED: udev rules pin each adapter by ID_PATH (physical USB port topology) to stable /dev/consoles/<name> symlinks that survive reboot regardless of enumeration order
  • conman owns serial devices directly (no ser2net/telnet in data path)
  • 7 log files active in /var/log/conman/
  • conmand enabled via systemd (survives reboot)
  • conmand systemd unit created (Debian package doesn't ship one)
  • ser2net installed but DISABLED (emergency TCP access only)
  • Old /root/conmap + manual screen workflow replaced
  • Stair-stepping fix: switched from ser2net→telnet→conman to conman direct-serial to eliminate telnet NVT CR stripping

Repo Merge

  • KNELServerBuild merged into PFVCluster (history preserved)
  • Directory structure reorganized (provisioning/, tests/, perf/, docs/)
  • All docs gardened: links fixed, stale refs removed, tailscale.md updated

Known Issues

Issue Impact Status
2 of 3 k3s cnodes on tsys4 tsys4 failure = etcd quorum lost (1 of 3) Accepted — fundamental constraint with 3 cnodes + 2 NFS servers. Revisit after Friday (local-SSD boot for one cnode). Deferred storage concern.
tsys1/5/9 RAM-overcommitted (alloc) 115%/134%/129% allocation; actual usage 73%/37%/62% — all under 80% goal Migrate NFS-backed VMs to tsys6/7 (117/171 GB free actual). See AUDIT-2026-07-29 §11
tsys5 is single-CPU CPU2 socket unpopulated; half the documented CPU capacity Consider seating a spare LGA771 Xeon (also activates CPU2 DIMM branch)
cnode4/5 stale VMs cnode4 (601) stopped; cnode5 (706) deleted cnode4 delete pending; cnode5 done ✓
Fleet CPU 0-5% idle Misses ~50%-idle utilization goal Consolidate baseline; solar-gate idle PowerEdge; grow via workloads
D2 IOPS hotspot 17 of 26 NFS-backed VMs on one spindle; S1/S4 underused Phase 2 disk migrations: D2→S2/S3/S4/D5 (see audit §11)
Technitium AXFR uses port 53 (occupied by Pi-hole) Zone transfer via rsync instead of native AXFR Workaround in place
tsys4/5 hardware pending Perf tuning incomplete on 2 hosts Friday physical install

Pending (next session priorities)

Full phased action checklist with current/target inventory is in proxmox/docs/AUDIT-2026-07-29.md §11. Summary below. SSD/NVMe tiering is deferred to post-Friday.

Phase 0 — Storage config fixes — DONE ✓ (tsys6 S1 typo, tsys9 S1 missing)

Phase 1 — Cleanup:

  • Delete cnode4 (601) on tsys6 (stopped, stale)
  • cnode5 (706) already deleted ✓

Phase 2 — Disk migrations off D2 (IOPS rebalancing, 6 moves):

  • wnode-tsys5 (509): D2→S4 · tsys-awx (600): D2→D5 · hfnoc-uisp (702): D2→S4
  • rr-middleware (703): D2→S4 · TCTC (704): D2→S3 · cnode2 (705): D2→S2

Phase 3 — Host + disk migrations (RAM/CPU relief, 6 VMs):

  • tsys1→tsys7: librenms(104), pdm(105), tailscale-router(109), kali-tsys(114), workbench(117)
  • tsys9→tsys6: tsys-siem(901)

Phase 4 — Grow wnode-tsys9: 4c/4GB → 6c/8GB (after siem leaves tsys9)

Phase 5 — Join wnodes to k3s (after migrations free resources)

Other:

  • Friday maintenance: tsys4 (PCIe NIC + RAM), tsys5 (2nd CPU candidate, NVMe, 2nd ethernet)
  • PDU: Rename outlets in powerman.conf (Friday onsite)
  • k8s deferred: ETL tooling, HPC scheduler, vcluster policy, solar-aware scale-out

Infrastructure Summary

Component Details
Proxmox hosts 7 standalone PVE, managed via PDM
DNS primary pfv-netinfra-01 (192.168.3.252) — Technitium + Pi-hole
DNS secondary pfv-netinfra-02 (192.168.3.253) — Technitium + Pi-hole
DNS production tailscale-router (read-only source of truth)
NTP pfv-netinfra-01/02 (redundant, LAN IPs, stratum 2/3)
Sandbox VM sectestbed-sandbox (VMID 6000 on pfv-tsys5)
Backup Proxmox Backup Server (PBS)
k8s control plane 3-node k3s HA (cnode1/2/3), all traffic over Tailscale
PDU Cyclades PM10i via powerman on pfv-tsys1 (port 10101)
Console 7 switch consoles via conman direct-serial on pfv-tsys4 (conmand:7890 on Tailscale)
Production Cloudron VPS, Reston VA (this cluster is R&D only)