diff --git a/STATUS.md b/STATUS.md index 346e14f..f730c25 100644 --- a/STATUS.md +++ b/STATUS.md @@ -90,35 +90,42 @@ infrastructure, Proxmox cluster ops, and k8s control plane. | Issue | Impact | Status | |-------|--------|--------| | **2 of 3 k3s cnodes on tsys4** | tsys4 failure = etcd quorum lost (1 of 3) | Accepted — fundamental constraint with 3 cnodes + 2 NFS servers. Revisit after Friday (local-SSD boot for one cnode). **Deferred storage concern.** | -| **tsys1/5/9 RAM-overcommitted (alloc)** | 115%/134%/129% allocation; **actual usage 72%/37%/60% — all under 80% goal** | Migrate NFS-backed VMs to tsys6/7 (117/171 GB free actual). See [AUDIT-2026-07-29](proxmox/docs/AUDIT-2026-07-29.md) §8 | +| **tsys1/5/9 RAM-overcommitted (alloc)** | 115%/134%/129% allocation; **actual usage 73%/37%/62% — all under 80% goal** | Migrate NFS-backed VMs to tsys6/7 (117/171 GB free actual). See [AUDIT-2026-07-29](proxmox/docs/AUDIT-2026-07-29.md) §11 | | **tsys5 is single-CPU** | CPU2 socket unpopulated; half the documented CPU capacity | Consider seating a spare LGA771 Xeon (also activates CPU2 DIMM branch) | | **cnode4/5 stale VMs** | cnode4 (601) stopped; cnode5 (706) deleted | cnode4 delete pending; cnode5 done ✓ | -| **Fleet CPU 0-6% idle** | Misses ~50%-idle utilization goal | Consolidate baseline; solar-gate idle PowerEdge; grow via workloads | +| **Fleet CPU 0-5% idle** | Misses ~50%-idle utilization goal | Consolidate baseline; solar-gate idle PowerEdge; grow via workloads | +| **D2 IOPS hotspot** | 17 of 26 NFS-backed VMs on one spindle; S1/S4 underused | Phase 2 disk migrations: D2→S2/S3/S4/D5 (see audit §11) | | Technitium AXFR uses port 53 (occupied by Pi-hole) | Zone transfer via rsync instead of native AXFR | Workaround in place | | tsys4/5 hardware pending | Perf tuning incomplete on 2 hosts | Friday physical install | ## Pending (next session priorities) -> **RAM/CPU placement actions** are itemized in -> [`proxmox/docs/AUDIT-2026-07-29.md`](proxmox/docs/AUDIT-2026-07-29.md) §12. -> Storage/disk actions are deferred to post-Friday (see below). +> **Full phased action checklist with current/target inventory is in** +> [`proxmox/docs/AUDIT-2026-07-29.md`](proxmox/docs/AUDIT-2026-07-29.md) §11. +> Summary below. SSD/NVMe tiering is deferred to post-Friday. -1. **VM migrations (RAM/CPU relief, zero hardware):** - - tsys1 → tsys7: librenms(104), pdm(105), workbench(117), tailscale-router(109), kali-tsys(114) - - tsys9 → tsys6: tsys-siem(901) -2. **Grow wnode-tsys9** to 6 vCPU / 16 GB (after siem move) — cross-compile node -3. **Stop + delete stale cnode4 (601) / cnode5 (706)** — not in the 3-node cluster -4. **Join wnodes to the k3s cluster** (after migrations free resources) -5. **Friday maintenance:** physical hardware + storage work (see - [`proxmox/docs/TODO.md`](proxmox/docs/TODO.md)) -6. **Wnode storage rebalancing** (disk — deferred): wnode-tsys6 D5→S4, wnode-tsys7 D5→S1 -7. **PDU:** Rename outlets in powerman.conf to match physical devices (Friday) -8. **k8s deferred topics:** ETL tooling, HPC scheduler, vcluster policy, - solar-aware scale-out +**Phase 0 — Storage config fixes — DONE ✓** (tsys6 S1 typo, tsys9 S1 missing) -> **See [`proxmox/docs/AUDIT-2026-07-29.md`](proxmox/docs/AUDIT-2026-07-29.md) -> for the consolidated audit with corrected storage topology, VM inventory, -> capacity analysis, and action checklist.** +**Phase 1 — Cleanup:** +- Delete cnode4 (601) on tsys6 (stopped, stale) +- cnode5 (706) already deleted ✓ + +**Phase 2 — Disk migrations off D2 (IOPS rebalancing, 6 moves):** +- wnode-tsys5 (509): D2→S4 · tsys-awx (600): D2→D5 · hfnoc-uisp (702): D2→S4 +- rr-middleware (703): D2→S4 · TCTC (704): D2→S3 · cnode2 (705): D2→S2 + +**Phase 3 — Host + disk migrations (RAM/CPU relief, 6 VMs):** +- tsys1→tsys7: librenms(104), pdm(105), tailscale-router(109), kali-tsys(114), workbench(117) +- tsys9→tsys6: tsys-siem(901) + +**Phase 4 — Grow wnode-tsys9:** 4c/4GB → 6c/8GB (after siem leaves tsys9) + +**Phase 5 — Join wnodes to k3s** (after migrations free resources) + +**Other:** +- **Friday maintenance:** tsys4 (PCIe NIC + RAM), tsys5 (2nd CPU candidate, NVMe, 2nd ethernet) +- **PDU:** Rename outlets in powerman.conf (Friday onsite) +- **k8s deferred:** ETL tooling, HPC scheduler, vcluster policy, solar-aware scale-out ## Infrastructure Summary diff --git a/proxmox/docs/AUDIT-2026-07-29.md b/proxmox/docs/AUDIT-2026-07-29.md index 40a7029..9f63776 100644 --- a/proxmox/docs/AUDIT-2026-07-29.md +++ b/proxmox/docs/AUDIT-2026-07-29.md @@ -187,7 +187,7 @@ These aren't causing problems but are underdocumented. Not actionable now. | 53100 | tsys-preprod-awx | 4 | 9 GB | local-nonprod (=S1) | LOCAL* | No | | 53101 | tsys-preprod-siem | 4 | 12 GB | local-nonprod (=S1) | LOCAL* | No | | 53102 | tsys-preprod-rancherplatform | 4 | 8 GB | local-nonprod (=S1) | LOCAL* | No | -| ×9 | sectestbed suite | 4 ea | 4 GB ea | local-nonprod (=S1) | LOCAL* | No | +| ×10 | sectestbed suite | 4 ea | 4 GB ea | local-nonprod (=S1) | LOCAL* | No | > *\*LOCAL = locked to tsys5. The local-nonprod VMs are on the S1 physical disk > but accessed via dir storage, not NFS — Proxmox will not migrate them. To