# K8S.md -- Kubernetes Architecture Deep-Dive **Date:** 2026-07-27 **Purpose:** Detailed kubernetes architecture plan for the pfv-k8s cluster. Companion to `PROJECT.md` (which has the fleet-wide assessment). **Status:** For discussion in a future session. No changes made. --- ## Table of Contents 1. [Workload Profile](#1-workload-profile) 2. [Current State](#2-current-state) 3. [Target Architecture](#3-target-architecture) 4. [Control Plane (Cnodes)](#4-control-plane-cnodes) 5. [Worker Nodes (Wnodes)](#5-worker-nodes-wnodes) 6. [Storage Class Design](#6-storage-class-design) 7. [ETL/HPC Considerations](#7-etlhpc-considerations) 8. [Migration Plan](#8-migration-plan) --- ## 1. Workload Profile This cluster runs **R&D and RackRental (containerlab) workloads** via Kubernetes. Production (Gitea, RustFS, Redmine, websites) lives on a VPS in Reston, VA running Cloudron. **Workload types expected:** | Type | Description | Storage need | RAM need | Examples | |------|------------|-------------|----------|---------| | **ETL (weather/GIS)** | Batch processing of large geospatial datasets. Sequential reads, transform, sequential writes. | High capacity (100s of GB), moderate IOPS | Medium (8-32 GB per job) | GRIB/NetCDF processing, raster reprojection | | **HPC (hardware startup)** | Compute-intensive simulations, firmware build pipelines, hardware-in-the-loop testing. | Low capacity, moderate IOPS | High (32-128 GB per job) | RTL simulation, PCB thermal analysis | | **RackRental (containerlab)** | Rapid deployment/teardown of network lab topologies. Many containers, short-lived. | Low capacity, high IOPS (container image pulls) | Low-Medium (4-16 GB) | Network topology testing, protocol validation | **Key storage insight:** ETL workloads need bulk capacity (NFS-HDD is fine -- sequential I/O). HPC and containerlab need low-latency random I/O (local SSD/NVMe is essential). The tiered StorageClass design (section 6) serves both. --- ## 2. Current State ### 2.1 pfv-k8s nodes and their storage | Node | Type | Host | Storage | Disk type | Status | |------|------|------|---------|-----------|--------| | cnode1 (107) | control | tsys1 | D5 (tsys4) | NFS-HDD | running | | cnode2 (603) | control | tsys6 | D2 (tsys4) | NFS-HDD | running | | cnode3 (106) | control | tsys1 | D2 (tsys4) | NFS-HDD | running | | cnode4 (601) | control | tsys6 | D2 (tsys4) | NFS-HDD | running | | cnode5 (602) | control | tsys6 | D5 (tsys4) | NFS-HDD | running | | wnode-tsys3 (313) | worker | tsys3 | D5 (tsys4) | NFS-HDD | running | | wnode-tsys5 (509) | worker | tsys5 | D2 (tsys4) | NFS-HDD | running | | wnode-tsys6 (100) | worker | tsys6 | D5 (tsys4) | NFS-HDD | **STOPPED** | | wnode-tsys7 (701) | worker | tsys7 | D5 (tsys4) | NFS-HDD | running | | wnode-tsys9 (905) | worker | tsys9 | S3 (tsys5) | NFS-HDD | running | ### 2.2 Problems 1. **100% of cnodes on tsys4 NFS.** D2 disk failure loses 3 of 5 cnodes = etcd quorum lost. 2. **90% of all k8s nodes on tsys4 NFS.** tsys4 failure kills the cluster. 3. **Zero nodes use SSD or NVMe.** All on NFS-over-HDD. 4. **Zero nodes use local-lvm.** tsys3/6/7/9 all have empty local storage (349 GB / 1.7 TB / 1.7 TB / 136 GB SSD respectively). 5. **wnode-tsys6 is stopped.** Reduces cluster capacity. 6. **3 cnodes on tsys6** -- should be on lighter hosts to free tsys6 for workers. --- ## 3. Target Architecture ### 3.1 Design principles 1. **Cnodes on lightweight hosts** (tsys1, tsys9, tsys3) -- frees tsys6/7 for heavy workers. 2. **Cnode storage split across tsys4 and tsys5** -- etcd survives either storage server failing. 3. **Wnode boot disks on local storage** -- eliminates NFS latency for container runtime and kubelet. 4. **Wnode data disks on NFS-HDD** -- bulk capacity for ETL/weather/GIS. 5. **tsys5 NVMe dedicated to wnode-tsys5** -- fastest tier for HPC jobs. 6. **One wnode per hypervisor host** -- maximize total cluster capacity. ### 3.2 Target node-host-storage matrix | Node | Type | Host | Boot disk | Data disk | Disk type | |------|------|------|-----------|-----------|-----------| | cnode1 | control | tsys1 | D5 (tsys4) | -- | NFS-HDD | | cnode2 | control | tsys9 | D2 (tsys4) | -- | NFS-HDD | | cnode3 | control | tsys1 | S2 (tsys5) | -- | NFS-HDD | | cnode4 | control | tsys9 | D5 (tsys4) | -- | NFS-HDD | | cnode5 | control | tsys3 | S3 (tsys5) | -- | NFS-HDD | | wnode-tsys1 | worker | tsys1 | D5 (tsys4) | -- | NFS-HDD (small) | | wnode-tsys2 | worker | tsys2 | **NVMe (960 PRO 512GB)** | **SATA SSD (850 EVO 1TB)** | **NVMe + SSD -- no NFS needed** | | wnode-tsys3 | worker | tsys3 | **local-lvm (NVMe PM961)** | S3 (NFS) | **LOCAL-NVMe** | | wnode-tsys5 | worker | tsys5 | **NVMe (local, Friday)** | local-nonprod (HDD) | **NVMe** | | wnode-tsys6 | worker | tsys6 | D2 (tsys4 NFS) | -- | NFS-HDD (local-lvm is USB 2.0 -- do not use) | | wnode-tsys7 | worker | tsys7 | D5 (tsys4 NFS) | -- | NFS-HDD (local-lvm is USB 2.0 -- do not use) | | wnode-tsys9 | worker | tsys9 | **local-lvm (SSD)** | S2 (NFS) | **LOCAL-SSD** | ### 3.3 Storage server distribution after changes | Storage server | cnodes | wnodes (boot) | wnodes (data) | |---------------|--------|---------------|---------------| | tsys4 (D2) | cnode2 | wnode-tsys6 | wnode-tsys7 | | tsys4 (D5) | cnode1, cnode4 | wnode-tsys1 | -- | | tsys5 (S2) | cnode3 | wnode-tsys9 | -- | | tsys5 (S3) | cnode5 | -- | wnode-tsys3 | **Note:** wnode-tsys2 needs no NFS (1.5 TB local SSD). wnode-tsys6/7 stay on NFS by design -- their local-lvm is USB 2.0 portable HDD (~30 MB/s), slower than NFS-HDD, and the user has chosen not to install internal drives. **No single disk or server is a quorum-losing failure point.** --- ## 4. Control Plane (Cnodes) ### 4.1 Cnode sizing Each cnode: 4 cores, 4 GB RAM, 32 GB disk. This is sufficient for etcd + kubernetes control plane components (API server, scheduler, controller-manager). ### 4.2 Cnode host placement rationale | Host | cnodes | RAM for cnodes | Total host RAM | Remaining for other VMs | |------|--------|---------------|---------------|------------------------| | tsys1 | 2 (cnode1, cnode3) | 8 GB | 32 GB | ~24 GB (but 11 infra VMs consume most) | | tsys9 | 2 (cnode2, cnode4) | 8 GB | 24 GB | ~16 GB (4 infra VMs + 1 wnode) | | tsys3 | 1 (cnode5) | 4 GB | 32 GB | ~28 GB (1 wnode at 20 GB = 8 GB headroom) | **tsys6 and tsys7 have ZERO cnodes** -- fully dedicated to heavy worker nodes. ### 4.3 Cnode storage placement rationale The 5 cnodes are split 3-on-tsys4 / 2-on-tsys5: | Disk | cnodes | Rationale | |------|--------|-----------| | D5 (tsys4 HDD) | cnode1, cnode4 | Spread load across 2 disks on tsys4 | | D2 (tsys4 HDD) | cnode2 | Only 1 cnode on D2 (was 3 -- reduces blast radius) | | S2 (tsys5 HDD) | cnode3 | tsys5 storage for quorum diversity | | S3 (tsys5 HDD) | cnode5 | tsys5 storage, different disk than S2 | **If D2 fails:** cnode2 dies. 4 of 5 survive. Quorum OK. **If D5 fails:** cnode1 + cnode4 die. 3 of 5 survive. Quorum OK. **If tsys4 fails:** cnode1, cnode2, cnode4 die. cnode3 + cnode5 survive on tsys5. **Only 2 of 5 -- QUORUM LOST.** Wait -- that is a problem. If tsys4 goes completely offline, we lose 3 cnodes and only have 2 on tsys5. That loses quorum (need 3). **Revision needed:** Move 1 more cnode to tsys5 storage. Target: 2 on tsys4, 3 on tsys5. But that means tsys5 failure (3 cnodes die) leaves only 2 on tsys4. Same problem inverted. The fundamental issue: with 5 cnodes and 2 storage servers, the best split is 3/2. The server holding 3 cnodes is a quorum-loss risk if it fails. The server holding 2 cnodes is safe (3 survive). **Proper solution: 3 cnodes on the "less likely to fail" server, 2 on the other.** After Friday's hardware work: - tsys4 will have a new PCIe NIC + 64 GB RAM -- more reliable - tsys5 will have bond0 fixed + NVMe -- more reliable Either way, 3/2 split means one server failure could lose quorum. **To truly solve this, use a 3rd storage target.** Options: - Use tsys9 local SSD for 1 cnode (breaks the 2-server model, adds a 3rd independent failure domain) - Use local-lvm on the cnode's own host (etcd data is local to the VM's host, no NFS dependency at all) **Best option: put cnode boot disks on local-lvm where available.** This eliminates NFS entirely for the control plane. Each cnode's etcd data lives on its own host's local disk -- no shared dependency. | cnode | Host | **Recommended storage** | Type | |-------|------|------------------------|------| | cnode1 | tsys1 | **local-lvm** (if space) or D5 (tsys4) | LOCAL-HDD or NFS-HDD | | cnode2 | tsys9 | **local-lvm (SSD)** | **LOCAL-SSD** | | cnode3 | tsys1 | **S2 (tsys5)** | NFS-HDD | | cnode4 | tsys9 | **local-lvm (SSD)** | **LOCAL-SSD** | | cnode5 | tsys3 | **local-lvm** | LOCAL-HDD | With this layout, a tsys4 failure takes down 0 cnodes. A tsys5 failure takes down 1 (cnode3). A host failure takes down at most 2 cnodes. Quorum always survives. **This is the recommended approach.** Local storage for cnodes wherever possible. NFS only as fallback. ### 4.4 etcd performance on local vs NFS | Storage | Typical fsync latency | etcd commit latency | Impact | |---------|----------------------|--------------------|--------| | NFS-HDD (via USB dongle on tsys4) | 5-15 ms | 10-30 ms | Slow API responses, sluggish pod scheduling | | NFS-HDD (via PCIe NIC, post-Friday) | 2-8 ms | 5-15 ms | Better but still network-bound | | Local HDD (tsys1/3/6/7 local-lvm) | 1-5 ms | 3-10 ms | No network hop, moderate improvement | | Local SSD (tsys9 PNY CS900) | 0.1-0.5 ms | 0.5-2 ms | **10-30x faster than NFS-HDD** | | NVMe (tsys5, Friday) | 0.02-0.1 ms | 0.1-0.5 ms | **100x faster than NFS-HDD** | etcd is the heartbeat of the kubernetes control plane. Every API call, every pod schedule, every controller reconciliation involves an etcd write. Cutting etcd commit latency from 15 ms to 1 ms makes the entire cluster feel 15x more responsive. **This is the single highest-impact change for k8s performance.** --- ## 5. Worker Nodes (Wnodes) ### 5.1 One wnode per hypervisor host | Host | wnode | Boot disk | Data disk | Total RAM | wnode RAM | Role | |------|-------|-----------|-----------|-----------|-----------|------| | tsys1 | wnode-tsys1 | D5 (tsys4 NFS) | -- | 32 GB | 4-8 GB | Small worker, infra co-tenant | | tsys2 | wnode-tsys2 | **NVMe (960 PRO 512GB)** | **SATA SSD (850 EVO 1TB)** | 32 GB | 16-24 GB | **Best storage of any worker -- 1.5TB local SSD, no NFS needed** | | tsys3 | wnode-tsys3 | **local-lvm (349 GB)** | S3 (NFS) | 32 GB | 20 GB | General worker | | tsys5 | wnode-tsys5 | **NVMe (local)** | local-nonprod (HDD) | 96 GB | 32-64 GB | **HPC/ETL powerhouse** | | tsys6 | wnode-tsys6 | D2 (tsys4 NFS) | -- | 128 GB | 64-96 GB | **Heavy worker, max RAM.** local-lvm is USB 2.0 -- stays on NFS | | tsys7 | wnode-tsys7 | D5 (tsys4 NFS) | -- | 192 GB | 96-128 GB | **Heavy worker, max RAM.** local-lvm is USB 2.0 -- stays on NFS | | tsys9 | wnode-tsys9 | **local-lvm SSD (136 GB)** | S2 (NFS) | 24 GB | 4-8 GB | Small worker, SSD boot | ### 5.2 Why boot disks on local-lvm Current: all wnodes boot from NFS. Every container image pull, every kubelet log write, every ephemeral volume traverses the NFS network path. With local-lvm boot disks: - **Container image pulls** write to local disk (100-150 MB/s HDD, no network hop) instead of NFS-HDD (80-120 MB/s with network latency) - **kubelet logs** stay local (no NFS writes for log rotation) - **ephemeral storage** (emptyDir volumes) uses local disk by default - **NFS server failure does not kill the wnode** -- the VM stays running, only the data disk (if mounted) goes away ### 5.3 Wnode sizing guidance | Host | Recommended wnode config | Rationale | |------|------------------------|-----------| | tsys7 (192 GB) | 8-12 cores, 96-128 GB RAM, NFS boot | Largest host -- run the heaviest ETL/HPC jobs here. local-lvm is USB 2.0 | | tsys6 (128 GB) | 8 cores, 64-96 GB RAM, NFS boot | Second-largest -- parallel heavy jobs. local-lvm is USB 2.0 | | tsys5 (96 GB + NVMe) | 4 cores, 32-64 GB RAM, NVMe boot + HDD data | NVMe makes this fastest for I/O-bound HPC | | tsys3 (32 GB) | 4 cores, 20 GB RAM, local-lvm boot | General-purpose worker | | tsys2 (32 GB, NVMe+SSD, incoming) | 4 cores, 16-24 GB RAM, **NVMe boot + SSD data** | **Fastest storage worker** -- HPC with I/O bounds | | tsys1 (32 GB) | 2 cores, 4-8 GB RAM | Small worker, don't starve infra VMs | | tsys9 (24 GB) | 2-4 cores, 4-8 GB RAM | Small worker, SSD boot is the advantage | ### 5.4 Tainting and labeling strategy Label wnodes by capability so the k8s scheduler can target them: ```yaml # Heavy RAM hosts (ETL/HPC) wnode-tsys6: workload=heavy, ram=128g wnode-tsys7: workload=heavy, ram=192g # NVMe host (I/O-intensive HPC) wnode-tsys5: workload=hpc, storage=nvme # SSD boot host (low-latency) wnode-tsys9: workload=light, storage=ssd # General workers wnode-tsys3: workload=general wnode-tsys2: workload=storage-fast, storage=nvme wnode-tsys1: workload=light ``` Then use nodeSelector or nodeAffinity in job specs: ```yaml # Weather/GIS ETL job -- needs lots of RAM spec: nodeSelector: workload: heavy # Firmware build -- needs fast storage spec: nodeSelector: storage: nvme ``` --- ## 6. Storage Class Design ### 6.1 Proposed StorageClasses | StorageClass | Provisioner | Where | Speed | Use case | |-------------|------------|-------|-------|----------| | `local-fast` | local-path (k8s) | wnode local-lvm / NVMe | 100-3500 MB/s | Container runtime, scratch, databases | | `nfs-hdd` | nfs-subdir-external-provisioner | tsys4 D2/D5, tsys5 S1-S4 | 80-120 MB/s | Bulk data, weather/GIS datasets | | `nfs-ssd` | nfs-subdir-external-provisioner | tsys4 D3, tsys5 T5-SSD | 200-400 MB/s | Latency-sensitive persistent data | ### 6.2 How this maps to wnode disk topology Each wnode has: - **Disk 1 (boot/OS):** local-lvm or NVMe. Contains the OS, kubelet, container runtime. k8s `local-fast` StorageClass provisioner points here. - **Disk 2 (bulk data, optional):** NFS mount. Mounted inside the VM as a second block device or filesystem. k8s `nfs-hdd` provisioner points here. Inside k8s, pods request storage via PVC: ```yaml # ETL job: needs bulk storage for weather data apiVersion: v1 kind: PersistentVolumeClaim spec: storageClassName: nfs-hdd accessModes: [ReadWriteMany] # NFS allows RWX resources: requests: storage: 500Gi # HPC job: needs fast scratch spec: storageClassName: local-fast accessModes: [ReadWriteOnce] resources: requests: storage: 50Gi ``` ### 6.3 NFS-SSD tier (D3 and T5-SSD -- both on tsys5 after Friday) The SSD NFS exports are a middle tier. **Both consolidated on tsys5** after the Friday D3 relocation: - **D3 (tsys5 SAS, 445 GB free):** CI/CD artifacts, container image cache, UCS/LDAP. Currently 99% empty. Fastest NFS tier after the move off USB. - **T5-SSD (tsys5 SAS, 140 GB free):** SIEM databases, monitoring data. Currently 35% used by ultix-streaming. **tsys5 is the fast-tier hub:** NVMe (local) + D3 SSD + T5-SSD all on one host. This simplifies the StorageClass design -- latency-sensitive PVCs target tsys5 exports, bulk PVCs target either server. ### 6.4 NFS data distribution across storage servers To avoid re-creating the "everything on tsys4" problem, distribute NFS data disks across both servers: | wnode | Boot (local) | Bulk data (NFS) | NFS server | |-------|-------------|-----------------|------------| | wnode-tsys3 | local-lvm | S3 | tsys5 | | wnode-tsys5 | NVMe | local-nonprod | local (no NFS) | | wnode-tsys6 | D2 (tsys4 NFS) | -- | tsys4 | | wnode-tsys7 | D5 (tsys4 NFS) | -- | tsys4 | | wnode-tsys9 | local-lvm (SSD) | S2 | tsys5 | This balances: 2 wnodes using tsys4 for bulk data, 2 using tsys5. --- ## 7. ETL/HPC Considerations ### 7.1 Weather/GIS ETL pipeline Typical flow: download GRIB/NetCDF files -> process (reproject, aggregate) -> store results. | Stage | Storage class | Why | |-------|-------------|-----| | Download raw data | `nfs-hdd` | Large sequential writes. NFS-HDD handles this well. | | Processing scratch | `local-fast` | Random access during transform. Local disk avoids NFS latency. | | Store results | `nfs-hdd` | Large sequential writes. Persistent. | **Recommendation:** Deploy a `local-fast` PV mount as `/scratch` on every wnode. ETL jobs use `/scratch` for intermediate processing and write final output to the NFS-mounted `/data`. ### 7.2 HPC workloads (hardware startup) Use cases: RTL simulation, PCB thermal analysis, firmware build pipelines. | Workload | Best wnode | Why | |----------|-----------|-----| | RTL simulation (CPU-bound, high RAM) | tsys7 (192 GB) | Most RAM, most cores (24t) | | Firmware builds (I/O-bound, moderate RAM) | tsys5 (NVMe) | Fastest storage for compile I/O | | Hardware-in-the-loop (latency-sensitive) | tsys9 (local SSD) | Lowest latency storage | | Parallel batch jobs | tsys6 + tsys7 | Distribute across both heavy hosts | ### 7.3 RackRental/containerlab Rapid container deployment. Key need: fast container image pulls. This is where **local-lvm boot disks** shine. Currently, every container image pull writes through NFS to a spinning disk -- slow. With local-lvm, images cache on local disk (even HDD is 2-3x faster than NFS-HDD for random I/O). On tsys9 (SSD) and tsys5 (NVMe), image pulls are near-instant. ### 7.4 Data locality for ETL For weather/GIS data that is read repeatedly (e.g., climate reanalysis), cache it on local-lvm of the heavy hosts: ``` tsys3 local-lvm (NVMe 349 GB): /data/cache/weather/ -- fastest cache tier tsys5 NVMe (local): /data/cache/gis/ -- fastest cache tier ``` **Note:** tsys6/7 local-lvm is USB 2.0 portable HDD (~30 MB/s) -- cannot be used for caching. Pre-populate weather/GIS data on D2/D5 (NFS) instead. This avoids re-reading the same data from the same NFS export on every job if the data is already cached in the page cache. --- ## 8. Migration Plan ### 8.1 Phase 1: Friday (after hardware work) After tsys5 cable + NVMe and tsys4 NIC + RAM: 1. **Format tsys5 NVMe** as local directory storage (e.g., `nvme-local`) 2. **Restart wnode-tsys6** (VM 100). Keep on NFS (D5). local-lvm is USB 2.0 -- do not use for VM storage. Recreate on D2 or D5 NFS. 3. **Move wnode-tsys9** (VM 905) disk from S3 (NFS) to local-lvm (SSD). ### 8.2 Phase 2: Cnode rebalance (maintenance window) These changes require creating new VMs on target hosts and migrating disks. Plan for a maintenance window with the k8s cluster briefly down. 1. Create cnode2 on tsys9 (local-lvm SSD if possible, or D2 NFS). 2. Create cnode4 on tsys9 (D5 NFS or local-lvm SSD). 3. Create cnode5 on tsys3 (S3 NFS or local-lvm). 4. Move cnode3 disk from D2 to S2 (tsys4 to tsys5). 5. Join new cnodes to etcd cluster, drain old cnode2/4/5, remove. ### 8.3 Phase 3: Wnode local storage migration (maintenance window) 1. Recreate wnode-tsys3 with boot disk on local-lvm (349 GB). 2. wnode-tsys6 stays on NFS (local-lvm is USB 2.0 HDD -- not suitable). 3. wnode-tsys7 stays on NFS (same reason). 4. Recreate wnode-tsys5 with boot disk on NVMe. 5. Add data disks (NFS) as second SCSI devices where applicable. ### 8.4 Phase 4: tsys2 integration (when rebuilt) 1. Install Proxmox on tsys2. 2. Run `scripts/check.sh` to inventory. 3. Run `scripts/apply-tunings.sh --apply`. 4. Create wnode-tsys2 with **boot disk on NVMe (960 PRO)** and **data disk on SATA SSD (850 EVO)**. No NFS needed -- 1.5 TB local SSD is the most local storage of any worker. 5. Join to k8s cluster. ### 8.5 Phase 5: Critical VM relocation 1. Move netinfra-02 (VM 904) from D2 to S3. 2. Move ucs-01 (VM 108) from D2 to D3 (SSD). 3. Move ucs-02 (VM 902) from D5 to tsys9 local-lvm (SSD). --- ## Open questions for next session 1. **Are the hosts a Proxmox cluster (pvecm) or standalone?** This determines whether live migration is available (huge simplification) or we need manual disk migration. Check `pvecm status` on each host. 2. **What k8s distribution is in use?** (k3s, kubeadm, RKE2?) This affects how nodes are joined/drain and how StorageClasses are configured. 3. **Container runtime?** (containerd, cri-o?) Affects local storage layout. 4. **Is there a container image registry mirror in the cluster?** Or do all pulls go to Docker Hub / external? A local registry on D3 SSD would speed up all pulls. 5. **What specific ETL tools?** (GDAL, PostGIS, xarray, Dask?) This affects whether jobs need shared (RWX) or exclusive (RWO) storage. 6. **HPC job scheduler?** (plain k8s Jobs, Argo Workflows, Volcano?) Affects how we label and taint nodes.