docs(audit): full ground-truth re-orientation + sync index files to module layout
Read-only audit (nmap/SNMP/SSH/API) reconciling the repo and fleet after the
operator's module/subproject restructure and infrastructure changes. New
authoritative audit doc captures: 4 switches up / 1 down (r2-tor-top offline),
63 VMs (was 28), 167 DNS zones (was 124), 753 DHCP leases (was 333), tsys5
storage now configured, tsys4 upgraded to 9.2.5, k8s cluster DOWN (cnode wiped,
no etcd quorum), netinfra-02 wiped, 3-tier Tailscale fleet (80 nodes).
Index files (README, AGENTS, docmap, STATUS) updated to the new layout
(provisioning→archive, powerman/console/ups→dcinfra, dns-cluster-setup/switches/
dhcp→netinfra, perf→proxmox/perf) and all stale root-level path references fixed.
💘 Generated with Crush
Assisted-by: Crush:glm-5.2
This commit is contained in:
@@ -3,19 +3,24 @@
|
||||
> **Human read-only. Agents maintain this file automatically after each work
|
||||
> session.** Do not edit by hand — the next agent run will overwrite it.
|
||||
|
||||
> **Last updated:** 2026-07-30 by Crush (GLM-5.2) — full ground-truth audit (network+storage+compute)
|
||||
> **Last updated:** 2026-08-05 by Crush (GLM-5.2) — full read-only ground-truth
|
||||
> re-orientation (nmap + SNMP + SSH + API queries).
|
||||
>
|
||||
> **Current authoritative audit:** [`proxmox/docs/AUDIT-2026-07-30.md`](proxmox/docs/AUDIT-2026-07-30.md)
|
||||
> — network topology, switch LAG configs, corrected storage topology, VM config
|
||||
> audit, fleet consistency, cross-rack NFS root-cause analysis.
|
||||
> **Current authoritative audit:** [`proxmox/docs/AUDIT-2026-08-05.md`](proxmox/docs/AUDIT-2026-08-05.md)
|
||||
> — repo restructure, 4 switches up / 1 down, 63 VMs, 167 DNS zones, 753 DHCP
|
||||
> leases, k8s cluster DOWN (cnode wiped), netinfra-02 wiped, tsys5 storage
|
||||
> configured, tsys4 upgraded, 3-tier Tailscale fleet (80 nodes).
|
||||
> Capacity targets are codified in
|
||||
> [`proxmox/docs/CAPACITY-GOALS.md`](proxmox/docs/CAPACITY-GOALS.md): 80% RAM
|
||||
> steady-state, ~50% idle CPU.
|
||||
> steady-state (fleet PASSES, highest tsys1 at 71%).
|
||||
|
||||
## Current State: STABLE
|
||||
## Current State: MIXED (infra stable; k8s + netinfra-02 need rebuild)
|
||||
|
||||
The merged PFVCluster repo is fully operational across provisioning, DNS
|
||||
infrastructure, Proxmox cluster ops, and k8s control plane.
|
||||
Provisioning, DNS primary (netinfra-01), Proxmox fleet, and monitoring stack
|
||||
are operational. **k8s control plane is DOWN** (cnode wiped, no etcd quorum)
|
||||
and **netinfra-02 is wiped** (needs rebuild as redundant secondary). Repo was
|
||||
restructured into a module/subproject layout (see
|
||||
[`proxmox/docs/AUDIT-2026-08-05.md`](proxmox/docs/AUDIT-2026-08-05.md)).
|
||||
|
||||
## Completed Work
|
||||
|
||||
@@ -34,7 +39,16 @@ infrastructure, Proxmox cluster ops, and k8s control plane.
|
||||
- [x] Zone replication via rsync systemd timer (every 60s)
|
||||
- [x] Credentials + 2FA replicated identically to production
|
||||
- [x] Both LAN IPs resolve knel.net device names + recurse externally
|
||||
- [x] **DHCP migrated from pfv-netboot → netinfra-01/02** (ISC failover-peer, primary/secondary). Both `normal`, 333 active leases load-balanced. DNS+NTP now handed out as 252/253. Netboot DHCP stopped + disabled. Webmin DHCP module on both nodes.
|
||||
- [x] **DHCP migrated from pfv-netboot → netinfra-01/02** (ISC failover-peer, primary/secondary). Both `normal`, **753 active leases** (was 333). DNS+NTP now handed out as 252/253. Netboot DHCP stopped + disabled. Webmin DHCP module on both nodes.
|
||||
- [ ] **netinfra-02 WIPED 2026-08-05** (storage move accident). Reachable on
|
||||
Tailscale (100.71.171.20) but SSH key auth fails — needs full rebuild as
|
||||
redundant secondary (key deploy + zone sync + DHCP failover-peer).
|
||||
- [x] **Technitium now runs as Docker container** (`tsys-dns`) on netinfra-01,
|
||||
not systemd dnsServer. Config at `/home/localuser/services/technitium/config`.
|
||||
**167 zones** (was 124). NTP = `ntpsec` (not chrony).
|
||||
- [x] **Pi-hole `/dev/shm` = 1.0 GB** (operator hotfix, was too small → prod
|
||||
outage). Healthcheck = `dig pi.hole` (DNS-based; gravity.db-specific check
|
||||
not yet added).
|
||||
|
||||
### Proxmox Cluster Ops
|
||||
- [x] 5 of 7 hosts fully performance-tuned (tsys1/3/6/7/9)
|
||||
@@ -53,8 +67,11 @@ infrastructure, Proxmox cluster ops, and k8s control plane.
|
||||
- [ ] tsys5: NVMe + SK hynix SSD installed but NOT configured
|
||||
- [ ] tsys2: pending rebuild from Win10 to Proxmox
|
||||
|
||||
### Kubernetes Control Plane (k3s HA — LIVE)
|
||||
- [x] 3-node k3s HA control plane deployed: cnode1/2/3 (v1.36.2+k3s1, embedded etcd)
|
||||
### Kubernetes Control Plane (k3s HA — DOWN, needs rebuild)
|
||||
- [ ] **Cluster DOWN 2026-08-05**: operator wiped a cnode + shut cnodes down
|
||||
(100% RAM on one). cnode2 has k3s API up, cnode1/cnode3 do not — no etcd
|
||||
quorum. Kubeconfig points to stale IP (100.97.178.106, matches no cnode).
|
||||
- [x] Was: 3-node k3s HA control plane (cnode1/2/3, embedded etcd, Tailscale)
|
||||
- [x] **All traffic over Tailscale IPs** — no LAN IPs in node status or certs
|
||||
- [x] All 3 cnodes tainted `control-plane:NoSchedule` (zero user workloads)
|
||||
- [x] 13/13 health checks pass (verify.sh): nodes Ready, etcd quorum,
|
||||
|
||||
Reference in New Issue
Block a user