Read-only audit (nmap/SNMP/SSH/API) reconciling the repo and fleet after the
operator's module/subproject restructure and infrastructure changes. New
authoritative audit doc captures: 4 switches up / 1 down (r2-tor-top offline),
63 VMs (was 28), 167 DNS zones (was 124), 753 DHCP leases (was 333), tsys5
storage now configured, tsys4 upgraded to 9.2.5, k8s cluster DOWN (cnode wiped,
no etcd quorum), netinfra-02 wiped, 3-tier Tailscale fleet (80 nodes).
Index files (README, AGENTS, docmap, STATUS) updated to the new layout
(provisioning→archive, powerman/console/ups→dcinfra, dns-cluster-setup/switches/
dhcp→netinfra, perf→proxmox/perf) and all stale root-level path references fixed.
💘 Generated with Crush
Assisted-by: Crush:glm-5.2
13 KiB
Project Status
Human read-only. Agents maintain this file automatically after each work session. Do not edit by hand — the next agent run will overwrite it.
Last updated: 2026-08-05 by Crush (GLM-5.2) — full read-only ground-truth re-orientation (nmap + SNMP + SSH + API queries).
Current authoritative audit:
proxmox/docs/AUDIT-2026-08-05.md— repo restructure, 4 switches up / 1 down, 63 VMs, 167 DNS zones, 753 DHCP leases, k8s cluster DOWN (cnode wiped), netinfra-02 wiped, tsys5 storage configured, tsys4 upgraded, 3-tier Tailscale fleet (80 nodes). Capacity targets are codified inproxmox/docs/CAPACITY-GOALS.md: 80% RAM steady-state (fleet PASSES, highest tsys1 at 71%).
Current State: MIXED (infra stable; k8s + netinfra-02 need rebuild)
Provisioning, DNS primary (netinfra-01), Proxmox fleet, and monitoring stack
are operational. k8s control plane is DOWN (cnode wiped, no etcd quorum)
and netinfra-02 is wiped (needs rebuild as redundant secondary). Repo was
restructured into a module/subproject layout (see
proxmox/docs/AUDIT-2026-08-05.md).
Completed Work
Server Provisioning (validated on sectestbed-sandbox)
- SetupNewSystem.sh deploys end-to-end to rc=0 (Debian 13 trixie)
- Security hardening: SSH, SCAP-STIG, Wazuh, 2FA (SSH+Cockpit+Webmin)
- NTP: redundant pfv-netinfra-01/02 (192.168.3.252/253), synced stratum 3
- DNS resolv.conf: managed static file pointing at netinfra pair
- Test suite: 5 tests (framework, safe-download, 2fa, https, system-req)
- VM validation harness: git-based deploy + auto-rollback + guest-agent access
DNS Cluster (pfv-netinfra-01/02)
- Production Technitium config replicated from tailscale-router (read-only)
- 124 zones on both nodes (knel.net + reverse DNS)
- pfv-netinfra-01 = PRIMARY, pfv-netinfra-02 = SECONDARY
- Zone replication via rsync systemd timer (every 60s)
- Credentials + 2FA replicated identically to production
- Both LAN IPs resolve knel.net device names + recurse externally
- DHCP migrated from pfv-netboot → netinfra-01/02 (ISC failover-peer, primary/secondary). Both
normal, 753 active leases (was 333). DNS+NTP now handed out as 252/253. Netboot DHCP stopped + disabled. Webmin DHCP module on both nodes. - netinfra-02 WIPED 2026-08-05 (storage move accident). Reachable on Tailscale (100.71.171.20) but SSH key auth fails — needs full rebuild as redundant secondary (key deploy + zone sync + DHCP failover-peer).
- Technitium now runs as Docker container (
tsys-dns) on netinfra-01, not systemd dnsServer. Config at/home/localuser/services/technitium/config. 167 zones (was 124). NTP =ntpsec(not chrony). - Pi-hole
/dev/shm= 1.0 GB (operator hotfix, was too small → prod outage). Healthcheck =dig pi.hole(DNS-based; gravity.db-specific check not yet added).
Proxmox Cluster Ops
- 5 of 7 hosts fully performance-tuned (tsys1/3/6/7/9)
- Full ground-truth audit 2026-07-30 (network + storage + compute + switches)
- Cross-rack NFS root cause identified: core-sw01 LAG hash=layer-2-3 (no L4), g16/g17 cable mismatch (3 of 4 links active), static mode (no LACP)
- tsys4/tsys5 storage changes documented: new NVMe (CT500P1 500GB), relocated SK hynix SSD, D3 changed to USB HDD, /mnt/albert + /mnt/backup added
- Console symlink bug fixed: udev rule singular→plural mismatch broke all 7 console ports after reboot
- Portable audit tooling built: probe-storage.sh, probe-network.sh, conman-console.py, snmp-switch-audit.py (AGPLv3-friendly, config-driven)
- Memory goal (≤80% actual): PASS fleet-wide — highest is tsys9 at 69%
- CPU ~50%-idle goal: fleet idles 0-5% (under-utilized; chase via workloads, not hardware)
- tsys4: PVE 9.1.5 (behind fleet 9.2.5), needs upgrade
- tsys5: NVMe + SK hynix SSD installed but NOT configured
- tsys2: pending rebuild from Win10 to Proxmox
Kubernetes Control Plane (k3s HA — DOWN, needs rebuild)
- Cluster DOWN 2026-08-05: operator wiped a cnode + shut cnodes down (100% RAM on one). cnode2 has k3s API up, cnode1/cnode3 do not — no etcd quorum. Kubeconfig points to stale IP (100.97.178.106, matches no cnode).
- Was: 3-node k3s HA control plane (cnode1/2/3, embedded etcd, Tailscale)
- All traffic over Tailscale IPs — no LAN IPs in node status or certs
- All 3 cnodes tainted
control-plane:NoSchedule(zero user workloads) - 13/13 health checks pass (verify.sh): nodes Ready, etcd quorum, Tailscale IPs, CoreDNS, API server, workload isolation
- Scripts in
k8s/: wipe, install-cp, join-servers, post-setup, verify - Kubeconfig saved to
~/.kube/config.pfv-k8s(gitignored, embedded certs) - Workers (wnodes) not yet joined to this cluster
- Distro decision: k3s chosen for regular R&D cluster. Talos docs
preserved in
k8s/docs/for future ITAR/classified cluster.
PDU Management (powerman on pfv-tsys1 — LIVE)
- Cyclades AlterPath PM10i (10 outlets) managed via powerman over serial
- USB-DB9 adapter (Prolific pl2303) with stable udev symlink
/dev/cyclades-pm10 - powermand listening on
127.0.0.1:10101+ Tailscale100.121.189.98:10101 - All 10 outlets defined as
outlet-1throughoutlet-10 - Validated: outlet 10 cycled off → on (8/8 test checks passed)
- Rename outlets to match physical devices (Friday onsite)
UPS Management (NUT on pfv-tsys1 — LIVE)
- APC Smart-UPS C 1500 fully reporting via NUT usbhid-ups + APC HID 0.100
- upsd listening on Tailscale
100.121.189.98:3493(port 3493) - Local upsmon (master) configured for graceful hypervisor shutdown
- Home Assistant credentials generated (
homeassistantuser in upsd.users) - Scripts in
ups/: discover, setup (idempotent), status - Home Assistant NUT integration configured via REST API (
ups/setup-ha-nut.sh) - Key fix: upsd now listens on LAN IP (192.168.3.11) in addition to Tailscale — HAOS runs Tailscale as an isolated add-on, so the HA container cannot route to Tailscale IPs. Shared vmbr0 bridge provides instant LAN path.
- Key fix: udev rule covers BOTH
usbandhidrawsubsystems (driver opens raw/dev/bus/usbafter setuid drop, not/dev/hidraw) - Tripp Lite UPS blocked: driver matches TrippLite HID 0.85 subdriver but fails reading 878-byte HID Report Descriptor (EAGAIN via USB hub). Needs physical reseat to direct motherboard USB port or cable replacement. Driver masked to prevent restart-loop spam.
- Home Assistant NUT integration: configured and live (3 sensors:
battery charge, status, status data). This APC (FW 02.2, mfg 2012) does not
report
ups.loadover USB HID — confirmed via both NUT and apcupsd tests. No wattage for Energy Dashboard; smart plug on UPS output is the practical path. apcupsd package conflicts with nut-server (mutually exclusive).
Console Management (ser2net telnet(rfc2217) + conman on pfv-tsys4 — LIVE)
- 7 network switch/router consoles managed via ser2net + conman
- USB enumeration problem SOLVED: udev rules pin each adapter by
ID_PATH (physical USB port topology) to stable
/dev/consoles/<name>symlinks that survive reboot regardless of enumeration order - ser2net exposes all 7 consoles on telnet(rfc2217) TCP ports (2001-2007)
bound to Tailscale IP only (
100.70.77.93) - conman connects to TCP ports for logging + multiplexing (7 log
files active in
/var/log/conman/) - Both ser2net + conmand enabled via systemd (survive reboot)
- conmand systemd unit created (Debian package doesn't ship one)
- conmand remote access enabled (loopback=off, port 7890)
- Old
/root/conmap+ manualscreenworkflow replaced - Stair-stepping fix: ser2net accepter changed from raw
tcptotelnet(rfc2217)so conman's telnet NVT negotiates binary mode properly - Validated 6/7 consoles working via conman from workstation (core-sw01, tor3-mgmt, tor3-stor, rrinfra-rtr, r2-tor-top, subodev-torsw)
- pfv-r2-sw (port 2007): DEAD — zero bytes at all baud rates (1200-115200), adapter present and mapped correctly (/dev/ttyUSB3). Physical issue: device powered off, console cable disconnected, or console port disabled in device config. Requires onsite investigation.
Ansible AWX (tsys-awx — LIVE)
- AWX 24.6.1 deployed via AWX Operator 2.19.1 on k3s single-node
- Accessible at
http://tsys-awx.knel.net(port 80, LoadBalancer via k3s ServiceLB) - Admin login:
admin/REDACTED_PASSWORD - PostgreSQL 15 (bundled, 8 GiB PVC on local-path)
- VM resized: 32→60 GB disk, swap converted to swapfile
- Scripts in
awx/: install-k3s, deploy-awx, verify-awx - HTTPS/TLS not yet configured (currently HTTP on port 80)
- Execution environments not yet customized
Repo Merge
- KNELServerBuild merged into PFVCluster (history preserved)
- Directory structure reorganized (provisioning/, tests/, perf/, docs/)
- All docs gardened: links fixed, stale refs removed, tailscale.md updated
Known Issues
| Issue | Impact | Status |
|---|---|---|
| core-sw01 LAG hash = layer-2-3 | NFS cross-rack capped at 1 Gbps per host pair | Fix: change to layer-2-3-4 (switch change, pending approval) |
| core-sw01 ch1 g16/g17 mismatch | Only 3 of 4 cross-rack links active | Fix: change member list g(13-15,17)→g(13-16) |
Cross-rack trunk uses mode on |
No LACP failure detection | Fix: change to mode active (both switches) |
| tsys5 bond0 degraded (1 of 2 slaves) | Single link, no redundancy | Physical: 2nd ethernet cable needed |
| tsys5 NVMe + SSD not configured | Best local storage unused | Config: mount + add to Proxmox storage.cfg |
| D2 carries 13 running VMs | IOPS hotspot persists | Migrate to S2/S3/new disks |
| tsys4 PVE 9.1.5 | Behind fleet (9.2.5) | Upgrade needed |
| tsys5 SSD 93% full | Samsung 860 PRO nearly exhausted | Offload to new SK hynix SSD |
| Port errors on core-sw01 | g9:2902, g38:590, g26:73 inErr | Investigate CRC/cable |
| tsys5 is single-CPU | CPU2 socket unpopulated | Consider seating spare LGA771 Xeon |
| Fleet CPU 0-5% idle | Misses ~50%-idle utilization goal | Consolidate baseline; grow via workloads |
Pending (next session priorities)
Full audit with findings + recommendations is in
proxmox/docs/AUDIT-2026-07-30.md.
Awaiting operator review of AUDIT-2026-07-30 before executing any fixes.
Switch changes (need approval):
- Fix core-sw01 ch1 member: g(13-15,17) → g(13-16) [restores 4th link]
- Fix core-sw01 LAG hash: layer-2-3 → layer-2-3-4 [unlocks multi-link NFS]
- Convert trunk to LACP: mode on → mode active [failure detection]
- Remove stale tsys8 config from tor3-stor (ch4/g13/g14)
Host-side changes (no switch access needed):
- Mount + configure tsys5 NVMe (CT500P1 500GB) as LVM-thin
- Mount + configure tsys5 SK hynix SSD (512GB, relocated from tsys4)
- Clean stale S3/S4 exports + storage.cfg on tsys5
- Upgrade tsys4 to PVE 9.2.5 / kernel 7.0.14
- Apply apply-tunings.sh to tsys4/5 (nconnect, bond hash, rmem/wmem)
- Audit VM cache modes (writethrough → writeback for NFS)
Other:
- Schedule: Operator onsite Monday 0700. Until then: software fixes, documentation, and OAM only (no physical access). Sunday off. Saturday = OAM Day (2026-08-01).
- Remote-capable (now–Saturday): tsys5 NVMe/SSD mount+config, fleet tuning (tsys4/5 nconnect/bond hash/rmem), VM config audit, OAM Day prep, documentation
- Onsite-only (Monday+): switch config changes (core-sw01 ch1 hash+member, tor3-stor LACP conversion), tsys5 2nd ethernet cable, physical hardware work
- Saturday OAM Day (2026-08-01): 100% instrumentation coverage — Beszel (metrics) on all hosts, Uptime Kuma (up/down) → Pushover alerts, LibreNMS (SNMP) for switches, Home Assistant (power/temp) for physical infra, centralized syslog (tool TBD)
- Investigate core-sw01 port errors (g9/g38/g26)
- k8s: join wnodes, label by workload class
- k8s deferred: ETL tooling, HPC scheduler, vcluster policy
Infrastructure Summary
| Component | Details |
|---|---|
| Proxmox hosts | 7 standalone PVE, managed via PDM (tsys8 retired) |
| Storage servers | tsys4 (NFS: D1/D2/D3/D5 + /mnt/albert + /mnt/backup), tsys5 (NFS: S1/S2 + local SSD + new NVMe unconfigured) |
| Switches | core-sw01 (PowerConnect 5448, rack5), tor3-stor/mgmt (PowerConnect 5324, rack3), + 1 unidentified (192.168.0.8) |
| DNS primary | pfv-netinfra-01 (192.168.3.252) — Technitium + Pi-hole |
| DNS secondary | pfv-netinfra-02 (192.168.3.253) — Technitium + Pi-hole |
| DNS production | tailscale-router (read-only source of truth) |
| NTP | pfv-netinfra-01/02 (redundant, LAN IPs, stratum 2/3) |
| Sandbox VM | sectestbed-sandbox (VMID 6000 on pfv-tsys5) |
| Backup | Proxmox Backup Server (PBS) |
| k8s control plane | 3-node k3s HA (cnode1/2/3), all traffic over Tailscale |
| PDU | Cyclades PM10i via powerman on pfv-tsys1 (port 10101) |
| AWX | tsys-awx: AWX 24.6.1 on k3s (http://tsys-awx.knel.net:80) |
| Console | 7 switch consoles via ser2net telnet(rfc2217)+conman on pfv-tsys4 (conmand:7890 on Tailscale) |
| UPS | APC Smart-UPS C 1500 via NUT on pfv-tsys1 (upsd:3493 on Tailscale) |
| Production | Cloudron VPS, Reston VA (this cluster is R&D only) |