docs: migrate all documentation to Discourse wiki topics
All knowledge docs (architecture, runbooks, references, audits, policies) have been migrated to community.turnsys.com as wiki topics in the VP TechOps category. Discourse is now the canonical source of truth for documentation; git edit history no longer serves as the doc changelog. 37 .md files converted to short pointers linking to their Discourse topics. AGENTS.md updated with new documentation workflow policy. Code (scripts, configs, playbooks) remains authoritative in git.
This commit is contained in:
@@ -1,452 +1,10 @@
|
||||
# Fleet Audit — 2026-07-29 (consolidated, storage-corrected)
|
||||
|
||||
> **The single authoritative audit.** Supersedes all prior audits and the
|
||||
> placement data in [`PROJECT.md`](PROJECT.md) §4-§8 where they conflict.
|
||||
> Capacity targets are codified in [`CAPACITY-GOALS.md`](CAPACITY-GOALS.md).
|
||||
|
||||
**Audit time:** 2026-07-29, fresh data from all 7 hosts (~05:09 CDT hardware
|
||||
dumps + ~05:40 CDT capacity probe). Storage configs read directly from
|
||||
`/etc/pve/storage.cfg`, `/etc/exports`, `/etc/fstab` on every host.
|
||||
**Hosts online:** tsys1, tsys3, tsys4, tsys5, tsys6, tsys7, tsys9
|
||||
**Hosts offline:** tsys2 (Win10, pending rebuild), tsys8 (retired / permanently offline)
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive summary
|
||||
|
||||
| Goal | Status | Detail |
|
||||
|------|--------|--------|
|
||||
| **Memory ≤ 80% actual steady-state** | **PASS (all 7 hosts)** | Highest is tsys1 at 73%. Fleet is healthy. |
|
||||
| **CPU ~50% at idle** | **MISS (all 7 hosts)** | Every host idles at **0-5% busy**. Under-utilized, not over-loaded. |
|
||||
|
||||
| Finding | Severity | § |
|
||||
|---------|----------|---|
|
||||
| **D2 carries 17 VMs** — the worst IOPS hotspot in the fleet | High | 5 |
|
||||
| **S1 is NOT empty — it carries 16 sectestbed VMs via `local-nonprod`** (previous audit was wrong) | High | 4, 5 |
|
||||
| ~~tsys6 S1 NFS export typo~~ (`S21`) — **FIXED by operator 2026-07-29**; all 8 stores active | Resolved | 6 |
|
||||
| ~~tsys9 missing S1~~ in storage.cfg — **FIXED by operator 2026-07-29**; S1 active | Resolved | 6 |
|
||||
| **`local-nonprod` on tsys5 = the S1 physical disk** (dir storage, not a separate disk) | Topology | 4 |
|
||||
| **tsys5 is single-CPU** (CPU2 socket unpopulated; prior "dual E5620" was wrong) | Medium | 9 |
|
||||
| **cnode4 (601) stopped, cnode5 (706) deleted** (user action this session) | Resolved | 10 |
|
||||
|
||||
---
|
||||
|
||||
## 2. Host capacity (live, 2026-07-29 ~05:40 CDT)
|
||||
|
||||
| Host | CPU (era) | Threads | RAM | Running VMs | Alloc RAM | Alloc% | **Actual%** | Swap used | CPU busy | Load |
|
||||
|------|-----------|---------|-----|-------------|-----------|--------|------------|-----------|----------|------|
|
||||
| tsys1 | i7-4770 Haswell (2013) | 8 | 31 GB | 11 | 36 GB | 115% | **73%** | 0 | 4% | 1.1 |
|
||||
| tsys3 | Xeon E3-1535M v5 Skylake (2015) | 8 | 31 GB | 1 | 28 GB | 90% | **10%** | 0 | 0% | 0.3 |
|
||||
| tsys4 | Xeon E3-1246 v3 Haswell (2013) | 8 | 16 GB | 1 | 2 GB | 13% | **14%** | 1.1 GB | 0% | 1.3 |
|
||||
| tsys5 | Xeon E5620 Westmere (2010) **single CPU** | 8 | 94 GB | 17 | 126 GB | 134% | **37%** | 0 | 5% | 4.0 |
|
||||
| tsys6 | 2× Xeon E5530 Nehalem (2009) | 16 | 126 GB | 2 | 44 GB | 35% | **4%** | 0 | 0% | 0.6 |
|
||||
| tsys7 | 2× Xeon E5-2630 v2 Ivy Bridge (2013) | 24 | 189 GB | 5 | 52 GB | 28% | **7%** | 0 | 1% | 0.9 |
|
||||
| tsys9 | i5-10500 Comet Lake (2020) | 12 | 23 GB | 6 | 30 GB | 129% | **62%** | 0 | 1% | 0.3 |
|
||||
|
||||
**Changes since first collection (~05:09):** cnode4 (601) stopped on tsys6,
|
||||
cnode5 (706) deleted from tsys7. No host migrations performed yet.
|
||||
|
||||
---
|
||||
|
||||
## 3. Storage topology (CORRECTED — read from live configs)
|
||||
|
||||
> **`storage.cfg` is per-host, not cluster-wide.** Each Proxmox node defines its
|
||||
> own storage entries. The same physical disk can appear under different storage
|
||||
> IDs on different hosts. This was the source of the prior audit's errors.
|
||||
|
||||
### 3.1 Physical disk → Proxmox storage ID mapping
|
||||
|
||||
| Physical disk | Server | NFS export | NFS ID (clients) | Local ID (on tsys5) | Type |
|
||||
|---------------|--------|------------|-----------------|---------------------|------|
|
||||
| WDC Red 3TB | tsys4 | `/mnt/tsys4/D2` | `D2` | — | HDD |
|
||||
| SK hynix SC300 512GB | tsys4 | `/mnt/tsys4/D3` | `D3` | — | SSD (USB) |
|
||||
| Hitachi 2TB | tsys4 | `/mnt/tsys4/D5` | `D5` | — | HDD |
|
||||
| *(disk1)* | tsys4 | `/mnt/tsys4/D1` | *(not in most storage.cfg)* | — | unknown |
|
||||
| *(disk7)* | tsys4 | `/mnt/tsys4/D7` | *(not in most storage.cfg)* | — | unknown |
|
||||
| **Seagate 916GB** | **tsys5** | **`/mnt/pfv-tsys5/S1`** | **`S1`** | **`local-nonprod`** | **HDD** |
|
||||
| Seagate 916GB | tsys5 | `/mnt/pfv-tsys5/S2` | `S2` | `S2` (NFS client) | HDD |
|
||||
| Seagate 916GB | tsys5 | `/mnt/pfv-tsys5/S3` | `S3` | `S3` (NFS client) | HDD |
|
||||
| Toshiba 458GB | tsys5 | `/mnt/pfv-tsys5/S4` | `S4` | `S4` (NFS client) | HDD |
|
||||
| **Samsung 860 PRO 234GB** | **tsys5** | **`/mnt/pfv-tsys5/ssd`** | **`T5-SSD`** | **`SSD`** (dir, local-only) | **SSD** |
|
||||
|
||||
### 3.2 Critical correction: `local-nonprod` = S1
|
||||
|
||||
On tsys5, the storage entry is:
|
||||
```
|
||||
dir: local-nonprod
|
||||
path /mnt/pfv-tsys5/S1
|
||||
```
|
||||
|
||||
**`local-nonprod` is NOT a separate disk.** It is the **S1 Seagate 916GB HDD**,
|
||||
accessed locally as a directory storage instead of via NFS. The 16 sectestbed
|
||||
and preprod VMs "on local-nonprod" are physically on S1. They appear locked to
|
||||
tsys5 because Proxmox treats `local-nonprod` (dir type) as a different storage
|
||||
than `S1` (NFS type), even though they point to the same directory.
|
||||
|
||||
**Implication:** S1 has **16 VMs of IOPS load**, not 0. The prior audit's plan
|
||||
to move production VMs onto S1 would have stacked them on top of 16 sandboxes.
|
||||
|
||||
### 3.3 Disk speed tiers (HDD only — SSD/NVMe deferred to Friday)
|
||||
|
||||
| Tier | Spindles | Speed | Current use |
|
||||
|------|----------|-------|-------------|
|
||||
| NFS-HDD (tsys4) | D2, D5 | 80-120 MB/s | Production VMs |
|
||||
| NFS-HDD (tsys5) | S1, S2, S3, S4 | 80-120 MB/s | Production + sectestbed VMs |
|
||||
| Local HDD (tsys5) | local-nonprod (=S1) | 80-120 MB/s | Sectestbed suite (locked to tsys5) |
|
||||
| SSD (tsys4 USB) | D3 | unreliable (USB 2.0) | Empty — do not use until Friday |
|
||||
| SSD (tsys5 local) | SSD dir (=Samsung 860 PRO) | 200-400 MB/s | ultix-streaming only |
|
||||
|
||||
---
|
||||
|
||||
## 4. Per-spindle IOPS load (CORRECTED)
|
||||
|
||||
| Spindle | Server | NFS-visible VMs | Local-only VMs | **Total physical load** |
|
||||
|---------|--------|----------------|----------------|------------------------|
|
||||
| **D2** (3TB) | tsys4 | **17** | 0 | **17 — worst hotspot** |
|
||||
| D5 (2TB) | tsys4 | 3 | 0 | 3 |
|
||||
| **S1** (916GB) | tsys5 | 0 | **16** (via local-nonprod) | **16 — second worst** |
|
||||
| S2 (916GB) | tsys5 | 4 | 0 | 4 |
|
||||
| S3 (916GB) | tsys5 | 3 | 0 | 3 |
|
||||
| S4 (458GB) | tsys5 | 1 | 0 | 1 — most spare capacity |
|
||||
|
||||
**D2 and S1 are both overloaded. S2/S3/S4 and D5 have room.**
|
||||
|
||||
The 16 sectestbed VMs on S1 are idle sandboxes — they don't generate much real
|
||||
IOPS, but they're all on one spindle. If any of them wake up simultaneously (CI
|
||||
pipeline, test suite), S1 becomes a bottleneck. **Do not add production VMs to
|
||||
S1.** Target S2, S3, S4, and D5 for IOPS spreading.
|
||||
|
||||
---
|
||||
|
||||
## 5. Storage bugs and config gaps found
|
||||
|
||||
### 5.1 BUG: tsys6 S1 export path typo (S21) — RESOLVED ✓
|
||||
|
||||
**Fixed by operator 2026-07-29.** The export path was corrected from `S21` to
|
||||
`S1`. Verified: `pvesm status` on tsys6 shows all 8 NFS stores active, including
|
||||
S1 (96 GB, 10.3% used by the sectestbed VMs visible via NFS).
|
||||
|
||||
### 5.2 GAP: tsys9 missing S1 in storage.cfg — RESOLVED ✓
|
||||
|
||||
**Fixed by operator 2026-07-29.** S1 entry added to tsys9's storage.cfg.
|
||||
Verified: all 8 NFS stores active on tsys9, S1 listing identical to tsys6.
|
||||
|
||||
### 5.3 ORPHANED: D1 and D7 exports on tsys4
|
||||
|
||||
tsys4 exports D1 and D7 via NFS, but neither appears in most hosts' storage.cfg.
|
||||
- **D1** holds KNELTemplate (4000) and RestoreTemplate (5500) — stopped
|
||||
templates only, no running VMs.
|
||||
- **D7** holds a disk for DellOpenManageEnterprise (500) — stopped.
|
||||
|
||||
These aren't causing problems but are underdocumented. Not actionable now.
|
||||
|
||||
---
|
||||
|
||||
## 6. Current VM inventory (host + disk, post-user-changes)
|
||||
|
||||
> cnode4 (601) = stopped. cnode5 (706) = deleted. All other VMs unchanged from
|
||||
> start of session. `Disk` = NFS export or local storage holding the primary
|
||||
> qcow2.
|
||||
|
||||
### tsys1 — Infrastructure (11 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server |
|
||||
|------|------|------|-----|------|--------|
|
||||
| 100 | pfv-bms (HomeAssistant) | 2 | 4 GB | D2 | tsys4 |
|
||||
| 101 | tsys-ca | 2 | 2 GB | D2 | tsys4 |
|
||||
| 102 | pfv-k8s-wnode-tsys1 | 4 | 4 GB | S3 | tsys5 |
|
||||
| 103 | **pfv-netinfra-01** | 2 | 2 GB | S2 | tsys5 |
|
||||
| 104 | tsys-librenms | 2 | 2 GB | D2 | tsys4 |
|
||||
| 105 | tsys-proxmox-datacenter | 2 | 2 GB | D2 | tsys4 |
|
||||
| 106 | pfv-k8s-cnode3 | 2 | 4 GB | D5 | tsys4 |
|
||||
| 108 | **tsys-ucs-01** | 4 | 8 GB | D2 | tsys4 |
|
||||
| 109 | tailscale-router | 2 | 2 GB | D2 | tsys4 |
|
||||
| 114 | kali-tsys | 2 | 2 GB | D2 | tsys4 |
|
||||
| 117 | tsys-secure-workbench | 2 | 4 GB | D2 | tsys4 |
|
||||
|
||||
### tsys3 — Compute (1 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server |
|
||||
|------|------|------|-----|------|--------|
|
||||
| 313 | pfv-k8s-wnode-tsys3 | 8 | 28 GB | D5 | tsys4 |
|
||||
|
||||
### tsys4 — Storage (1 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server |
|
||||
|------|------|------|-----|------|--------|
|
||||
| 400 | pfv-proxmox-backup-server | 2 | 2 GB | local-lvm | LOCAL |
|
||||
|
||||
### tsys5 — Storage + sandboxes (17 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server | Migratable? |
|
||||
|------|------|------|-----|------|--------|-------------|
|
||||
| 509 | pfv-k8s-wnode-tsys5 | 8 | 32 GB | D2 | tsys4 | Yes (NFS) |
|
||||
| 5111 | ultix-streaming | 4 | 9 GB | SSD (local dir) | LOCAL | No (local) |
|
||||
| 5112 | ultix-offstage | 4 | 6 GB | local-lvm | LOCAL | No (local) |
|
||||
| 5101 | sectestbed-siem | 4 | 10 GB | local-nonprod (=S1) | LOCAL* | No (dir storage) |
|
||||
| 53100 | tsys-preprod-awx | 4 | 9 GB | local-nonprod (=S1) | LOCAL* | No |
|
||||
| 53101 | tsys-preprod-siem | 4 | 12 GB | local-nonprod (=S1) | LOCAL* | No |
|
||||
| 53102 | tsys-preprod-rancherplatform | 4 | 8 GB | local-nonprod (=S1) | LOCAL* | No |
|
||||
| ×10 | sectestbed suite | 4 ea | 4 GB ea | local-nonprod (=S1) | LOCAL* | No |
|
||||
|
||||
> *\*LOCAL = locked to tsys5. The local-nonprod VMs are on the S1 physical disk
|
||||
> but accessed via dir storage, not NFS — Proxmox will not migrate them. To
|
||||
> unlock, move disk from `local-nonprod` → `S1` (same physical disk, changes the
|
||||
> storage ID to the NFS-shared one). Low priority — these are idle sandboxes.*
|
||||
|
||||
### tsys6 — RackRental-ready (2 running, 1 stopped)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server | Status |
|
||||
|------|------|------|-----|------|--------|--------|
|
||||
| 100 | pfv-k8s-wnode-tsys6 | 4 | 32 GB | S4 | tsys5 | running |
|
||||
| 600 | tsys-awx | 4 | 12 GB | D2 | tsys4 | running |
|
||||
| 601 | pfv-k8s-cnode4 | 4 | 4 GB | D2 | tsys4 | **STOPPED** |
|
||||
|
||||
### tsys7 — RackRental-ready (5 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server |
|
||||
|------|------|------|-----|------|--------|
|
||||
| 701 | pfv-k8s-wnode-tsys7 | 4 | 32 GB | S3 | tsys5 |
|
||||
| 702 | hfnoc-uisp | 4 | 8 GB | D2 | tsys4 |
|
||||
| 703 | rr-middleware | 2 | 2 GB | D2 | tsys4 |
|
||||
| 704 | TCTC | 4 | 6 GB | D2 | tsys4 |
|
||||
| 705 | pfv-k8s-cnode2 | 4 | 4 GB | D2 | tsys4 |
|
||||
|
||||
### tsys9 — Infra + control plane (6 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server |
|
||||
|------|------|------|-----|------|--------|
|
||||
| 901 | tsys-siem | 2 | 8 GB | D2 | tsys4 |
|
||||
| 902 | **tsys-ucs-02** | 4 | 8 GB | S2 | tsys5 |
|
||||
| 903 | kali-rd | 2 | 2 GB | D5 | tsys4 |
|
||||
| 904 | **pfv-netinfra-02** | 2 | 4 GB | D2 | tsys4 |
|
||||
| 905 | pfv-k8s-wnode-tsys9 | 4 | 4 GB | S2 | tsys5 |
|
||||
| 906 | pfv-k8s-cnode1 | 2 | 4 GB | S3 | tsys5 |
|
||||
|
||||
### k3s control plane (ground truth via `kubectl`)
|
||||
|
||||
3 cnodes joined: cnode1 (906/tsys9), cnode2 (705/tsys7), cnode3 (106/tsys1).
|
||||
All control-plane,etcd role. CPU 2%, mem ~1.1 GB each. **No wnodes joined yet.**
|
||||
|
||||
---
|
||||
|
||||
## 7. Workload placement model
|
||||
|
||||
| Workload family | Needs | Best hosts | Why |
|
||||
|-----------------|-------|-----------|-----|
|
||||
| **ETL / HPC / cross-compile** | High IPC + fast storage | **tsys9, tsys3, tsys2(in)** | Newest CPUs + local SSD/NVMe |
|
||||
| **RackRental.net** | Massive RAM, CPU-light | **tsys7, tsys6** | 189/126 GB RAM; weak CPU fine for idle labs |
|
||||
| **Infra + k8s control plane** | Always-on, low resource | **tsys1, tsys9** | Efficient enough; cnodes spread for etcd diversity |
|
||||
|
||||
---
|
||||
|
||||
## 8. Answers to operator's questions
|
||||
|
||||
### 8.1 "Should I make the k8s wnode VMs bigger on tsys1/9 for ETL/HPC?"
|
||||
|
||||
**tsys9 — YES.** Newest CPU (i5-10500, 12 threads) + local SSD. Currently 4c/4GB.
|
||||
After moving tsys-siem (8 GB) to tsys6, grow to **6c/8-16 GB**.
|
||||
|
||||
**tsys1 — NO.** The i7-4770 is Haswell-2013 (2nd-oldest). tsys1 is the most
|
||||
RAM-loaded host (73% actual, 115% alloc). It should shed VMs, not grow its wnode.
|
||||
|
||||
### 8.2 "Are tsys6/7 better for RackRental.net?"
|
||||
|
||||
**Yes.** RAM-bound, CPU-light labs fit 126/189 GB perfectly. They don't need
|
||||
more RAM (4%/7% actual). Their bottleneck is disk (USB 2.0 local-lvm ≈ 30 MB/s)
|
||||
— deferred to Friday + StorageClass work.
|
||||
|
||||
### 8.3 "How should I deploy the DDR3?"
|
||||
|
||||
Verify type first. tsys5/6/7 require **RDIMM ECC** (consumer UDIMM won't POST).
|
||||
tsys1 is maxed (32 GB, caps at 8 GB/slot). tsys4 caps at 8 GB/slot (max 32 GB).
|
||||
Most likely only useful for tsys4 16→32 GB. The real RAM lever is **free VM
|
||||
migration**, not sticks. See [`CAPACITY-GOALS.md`](CAPACITY-GOALS.md) §5.
|
||||
|
||||
---
|
||||
|
||||
## 9. Hardware corrections
|
||||
|
||||
**tsys5 is single-CPU.** dmidecode: `CPU2: Status Unpopulated`. The T7500 has
|
||||
one E5620 (4c/8t), not two. The 6 empty DIMM slots are on the CPU2 branch —
|
||||
useless without a 2nd CPU. Seating a spare LGA771 Xeon would double CPU capacity
|
||||
and activate the CPU2 DIMM branch.
|
||||
|
||||
**tsys4 is using ~1.1 GB of swap.** Only host with active swap. Minor, but
|
||||
monitor after the planned RAM upgrade.
|
||||
|
||||
---
|
||||
|
||||
## 10. Target end-state inventory (RAM/CPU + HDD IOPS balanced)
|
||||
|
||||
> All disks stay on **HDD** (SSD/NVMe deferred to Friday). Goals: (a) relieve
|
||||
> tsys1/9 RAM over-allocation, (b) relieve D2 (17→5 VMs), (c) spread to S2/S3/S4
|
||||
> (currently 4/3/1), (d) **avoid S1** (already has 16 local VMs), (e) keep HA
|
||||
> pairs split across storage servers.
|
||||
|
||||
### tsys1 — Infrastructure (6 running, ~24 GB alloc → 77%)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server | Change |
|
||||
|------|------|------|-----|------|--------|--------|
|
||||
| 100 | pfv-bms | 2 | 4 GB | D2 | tsys4 | none |
|
||||
| 101 | tsys-ca | 2 | 2 GB | D2 | tsys4 | none |
|
||||
| 102 | pfv-k8s-wnode-tsys1 | 4 | 4 GB | S3 | tsys5 | none |
|
||||
| 103 | **pfv-netinfra-01** | 2 | 2 GB | S2 | tsys5 | none |
|
||||
| 106 | pfv-k8s-cnode3 | 2 | 4 GB | D5 | tsys4 | none |
|
||||
| 108 | **tsys-ucs-01** | 4 | 8 GB | D2 | tsys4 | none |
|
||||
|
||||
### tsys3 — Compute (1 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server | Change |
|
||||
|------|------|------|-----|------|--------|--------|
|
||||
| 313 | pfv-k8s-wnode-tsys3 | 8 | 28 GB | D5 | tsys4 | none |
|
||||
|
||||
### tsys4 — Storage (1 running)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Change |
|
||||
|------|------|------|-----|------|--------|
|
||||
| 400 | pfv-proxmox-backup-server | 2 | 2 GB | local-lvm | none |
|
||||
|
||||
### tsys5 — Storage + sandboxes (17 running, unchanged)
|
||||
|
||||
All 17 stay. wnode-tsys5 (509) disk moves D2→S4. The 16 local-nonprod VMs stay
|
||||
locked (idle sandboxes, low priority).
|
||||
|
||||
### tsys6 — RackRental-ready (3 running, ~52 GB alloc → 41%)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server | Change |
|
||||
|------|------|------|-----|------|--------|--------|
|
||||
| 100 | pfv-k8s-wnode-tsys6 | 4 | 32 GB | S4 | tsys5 | none |
|
||||
| 600 | tsys-awx | 4 | 12 GB | **D5** | tsys4 | **disk D2→D5** |
|
||||
| 901 | tsys-siem | 2 | 8 GB | **S2** | tsys5 | **host tsys9→tsys6, disk D2→S2** |
|
||||
| ~~601~~ | ~~cnode4~~ | — | — | — | — | **DELETE (stopped, stale)** |
|
||||
|
||||
### tsys7 — RackRental-ready (10 running, ~64 GB alloc → 34%)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server | Change |
|
||||
|------|------|------|-----|------|--------|--------|
|
||||
| 701 | pfv-k8s-wnode-tsys7 | 4 | 32 GB | S3 | tsys5 | none |
|
||||
| 702 | hfnoc-uisp | 4 | 8 GB | **S4** | tsys5 | **disk D2→S4** |
|
||||
| 703 | rr-middleware | 2 | 2 GB | **S4** | tsys5 | **disk D2→S4** |
|
||||
| 704 | TCTC | 4 | 6 GB | **S3** | tsys5 | **disk D2→S3** |
|
||||
| 705 | pfv-k8s-cnode2 | 4 | 4 GB | **S2** | tsys5 | **disk D2→S2** |
|
||||
| 104 | tsys-librenms | 2 | 2 GB | **S3** | tsys5 | **host tsys1→tsys7, disk D2→S3** |
|
||||
| 105 | tsys-proxmox-datacenter | 2 | 2 GB | **S4** | tsys5 | **host tsys1→tsys7, disk D2→S4** |
|
||||
| 109 | tailscale-router | 2 | 2 GB | D2 | tsys4 | **host tsys1→tsys7** (disk stays D2) |
|
||||
| 114 | kali-tsys | 2 | 2 GB | **S2** | tsys5 | **host tsys1→tsys7, disk D2→S2** |
|
||||
| 117 | tsys-secure-workbench | 2 | 4 GB | **S3** | tsys5 | **host tsys1→tsys7, disk D2→S3** |
|
||||
|
||||
### tsys9 — Infra + control + compute (5 running, ~22 GB alloc → 94%)
|
||||
|
||||
| VMID | Name | vCPU | RAM | Disk | Server | Change |
|
||||
|------|------|------|-----|------|--------|--------|
|
||||
| 902 | **tsys-ucs-02** | 4 | 8 GB | S2 | tsys5 | none |
|
||||
| 903 | kali-rd | 2 | 2 GB | D5 | tsys4 | none |
|
||||
| 904 | **pfv-netinfra-02** | 2 | 4 GB | D2 | tsys4 | none |
|
||||
| 905 | pfv-k8s-wnode-tsys9 | **6** | **8 GB** | S2 | tsys5 | **grow 4c/4G→6c/8G** |
|
||||
| 906 | pfv-k8s-cnode1 | 2 | 4 GB | S3 | tsys5 | none |
|
||||
| ~~901~~ | ~~tsys-siem~~ | — | — | — | — | **move→tsys6** |
|
||||
|
||||
### End-state per-spindle IOPS load
|
||||
|
||||
| Spindle | Server | # VMs | VMs on it |
|
||||
|---------|--------|-------|-----------|
|
||||
| **D2** (3TB) | tsys4 | **5** | bms, ca, ucs01, tailscale, netinfra02 |
|
||||
| D5 (2TB) | tsys4 | 4 | cnode3, wnode3, awx, kali-rd |
|
||||
| S1 (916GB) | tsys5 | **16** (unchanged) | sectestbed suite (local-nonprod, idle) |
|
||||
| **S2** (916GB) | tsys5 | **7** | netinfra01, ucs02, wnode9, cnode2, siem, kali, pdm* |
|
||||
| **S3** (916GB) | tsys5 | **7** | wnode1, wnode7, cnode1, uisp, TCTC, librenms, workbench |
|
||||
| **S4** (458GB) | tsys5 | **5** | wnode5, wnode6, rr-middleware, uisp*... |
|
||||
|
||||
> *S2/S3 carry 7 each because the sectestbed load on S1 is idle. If those 16
|
||||
> sectestbed VMs ever become active, rebalance — but for now S2/S3 at 7
|
||||
> production VMs each is fine (they're 916 GB disks with light per-VM I/O).*
|
||||
|
||||
**D2: 17→5. No VMs added to S1. S2/S3/S4 absorb the load.**
|
||||
|
||||
HA pairs stay split:
|
||||
- netinfra: 01 on S2 (tsys5), 02 on D2 (tsys4) ✓
|
||||
- ucs: 01 on D2 (tsys4), 02 on S2 (tsys5) ✓
|
||||
|
||||
---
|
||||
|
||||
## 11. Action checklist (do in this order)
|
||||
|
||||
### Phase 0 — fix storage config bugs first — DONE ✓
|
||||
|
||||
- ~~Fix tsys6 S1 typo~~ — **operator fixed 2026-07-29** (verified: 8/8 stores active)
|
||||
- ~~Add S1 to tsys9 storage.cfg~~ — **operator fixed 2026-07-29** (verified)
|
||||
|
||||
### Phase 1 — cleanup
|
||||
|
||||
- **Delete cnode4 (VMID 601) on tsys6** — stopped, stale, not in cluster
|
||||
- **Confirm cnode5 (706) gone** — already deleted by operator ✓
|
||||
|
||||
### Phase 2 — disk migrations off D2 (IOPS rebalancing)
|
||||
|
||||
> Do disk moves on VMs that are staying on their current host first.
|
||||
|
||||
- **wnode-tsys5 (509):** disk **D2 → S4** (host stays tsys5)
|
||||
- **tsys-awx (600):** disk **D2 → D5** (host stays tsys6)
|
||||
- **hfnoc-uisp (702):** disk **D2 → S4** (host stays tsys7)
|
||||
- **rr-middleware (703):** disk **D2 → S4** (host stays tsys7)
|
||||
- **TCTC (704):** disk **D2 → S3** (host stays tsys7)
|
||||
- **cnode2 (705):** disk **D2 → S2** (host stays tsys7)
|
||||
|
||||
### Phase 3 — host + disk migrations (RAM/CPU relief, 6 VMs)
|
||||
|
||||
- **tsys-librenms (104):** host **tsys1 → tsys7**, disk **D2 → S3**
|
||||
- **tsys-proxmox-datacenter (105):** host **tsys1 → tsys7**, disk **D2 → S4**
|
||||
- **tailscale-router (109):** host **tsys1 → tsys7** (disk stays **D2**)
|
||||
- **kali-tsys (114):** host **tsys1 → tsys7**, disk **D2 → S2**
|
||||
- **tsys-secure-workbench (117):** host **tsys1 → tsys7**, disk **D2 → S3**
|
||||
- **tsys-siem (901):** host **tsys9 → tsys6**, disk **D2 → S2**
|
||||
|
||||
### Phase 4 — grow the compute wnode
|
||||
|
||||
- **wnode-tsys9 (905):** resize **4c/4GB → 6c/8GB** (stop/start required)
|
||||
|
||||
### Phase 5 — join workers to k3s
|
||||
|
||||
- Join wnodes via `k8s/join-servers.sh`
|
||||
- Label: `workload=compute` (tsys3/9), `workload=rackrental` (tsys6/7), `workload=light` (tsys1/5)
|
||||
|
||||
### Verification after each phase
|
||||
|
||||
- `qm list` per host — confirm VM count + status
|
||||
- `cat /proc/meminfo` — confirm actual RAM ≤ 80%
|
||||
- `kubectl top nodes` — confirm cnodes healthy
|
||||
- `ssh root@pfv-tsys4 'df -h /mnt/tsys4/D2'` — confirm D2 frees up
|
||||
- `ssh root@pfv-tsys5 'df -h /mnt/pfv-tsys5/S2 /mnt/pfv-tsys5/S3 /mnt/pfv-tsys5/S4'` — confirm fills
|
||||
|
||||
---
|
||||
|
||||
## 12. Future: local storage as ephemeral scratch
|
||||
|
||||
> Codified as standing policy in [`CAPACITY-GOALS.md`](CAPACITY-GOALS.md) §6.
|
||||
|
||||
The fleet has substantial **unused local storage IOPS** on several hosts. OS
|
||||
drives stay on NFS (no per-node disk redundancy — see policy), but local storage
|
||||
is ideal for **ephemeral scratch** (container image cache, build artifacts,
|
||||
`emptyDir` volumes) where data loss on node failure is acceptable.
|
||||
|
||||
**Per-host opportunity:**
|
||||
|
||||
| Host | Local | Available | Scratch-suitable? |
|
||||
|------|-------|-----------|-------------------|
|
||||
| tsys3 | NVMe (PM961) | 349 GB | **Yes — best existing** |
|
||||
| tsys9 | SSD (PNY CS900) | 136 GB | **Yes** |
|
||||
| tsys2 | NVMe + SSD | 512 GB + 1 TB | **Yes (incoming)** |
|
||||
| tsys5 | HDD (Hitachi) | ~1.7 TB | Yes (bulk) |
|
||||
| tsys6 | USB 2.0 HDD | 1.7 TB | **NO (30 MB/s)** |
|
||||
| tsys7 | USB 2.0 HDD | 1.7 TB | **NO (30 MB/s)** |
|
||||
|
||||
**The tsys6/7 problem:** they are the RackRental hosts (most in need of fast
|
||||
image pulls) but their local storage is USB 2.0 — slower than NFS. The real fix
|
||||
is installing internal drives (tsys6 has a disabled SAS controller; tsys7 has 5
|
||||
empty SATA ports). Deferred to a physical-hardware session.
|
||||
|
||||
**Implementation:** post-Friday, once SSD/NVMe tiers are installed and the k8s
|
||||
StorageClass design is finalized. Wnodes with local storage get a `/scratch`
|
||||
mount for ephemeral data; persistent data stays on NFS. See
|
||||
[`CAPACITY-GOALS.md`](CAPACITY-GOALS.md) §6.4 for the StorageClass mapping.
|
||||
# proxmox/docs/AUDIT-2026-07-29.md
|
||||
|
||||
> **Documentation moved to Discourse — the canonical source of truth.**
|
||||
>
|
||||
> **Audit snapshot 2026-07-29**
|
||||
>
|
||||
> **Read it here:** https://community.turnsys.com/t/298
|
||||
>
|
||||
> *Migrated 2026-08-06. This file is kept as a pointer for git-browsing context.
|
||||
> Do not update content here — edit the Discourse wiki topic instead.*
|
||||
|
||||
Reference in New Issue
Block a user