docs: migrate all documentation to Discourse wiki topics

All knowledge docs (architecture, runbooks, references, audits, policies)
have been migrated to community.turnsys.com as wiki topics in the VP
TechOps category. Discourse is now the canonical source of truth for
documentation; git edit history no longer serves as the doc changelog.

37 .md files converted to short pointers linking to their Discourse
topics. AGENTS.md updated with new documentation workflow policy.
Code (scripts, configs, playbooks) remains authoritative in git.
This commit is contained in:
2026-08-06 14:08:34 -05:00
parent d6988bb35f
commit 88faf35ec3
37 changed files with 378 additions and 8239 deletions
+9 -124
View File
@@ -1,125 +1,10 @@
# DHCP Migration: pfv-netboot → pfv-netinfra-01/02
# netinfra/dhcp-migration.md
**Date:** 2026-07-29
**Status:****COMPLETE** — both servers in `normal` failover state, 333
active leases load-balanced. Netboot DHCP stopped + disabled.
> **Safety nets:** pfv-netboot, pfv-netinfra-01/02, sectestbed-sandbox, and
> ultix-streaming all have static IPs (reachable via Tailscale even if DHCP
> fails). sectestbed-sandbox snapshot `pre-dhcp-migration` exists as rollback.
## Migration results (2026-07-29 17:20 CDT)
| Check | Result |
|-------|--------|
| netinfra-01 (primary) | `active`, failover `normal`, 244 active leases |
| netinfra-02 (secondary) | `active`, failover `normal`, 89 active leases |
| Failover partnership | "Both servers normal" — load-balanced at split=128 |
| DHCP test (sandbox) | Lease obtained from 192.168.3.252, IP 192.168.3.50/22 |
| DNS handed out | `192.168.3.252, 192.168.3.253` (netinfra pair) |
| NTP handed out | `192.168.3.252, 192.168.3.253` (netinfra pair) |
| Domain | `knel.net` |
| Gateway | `192.168.3.254` |
| Webmin | Port 10000 (SSL) on both nodes; DHCP module configured |
| Netboot DHCP | Stopped + disabled (won't start on reboot) |
| Netmask fix | Both nodes corrected /24 → /22 |
## 1. Current state (discovered 2026-07-29)
| | pfv-netboot (SOURCE) | pfv-netinfra-01 (TARGET primary) | pfv-netinfra-02 (TARGET secondary) |
|---|---|---|---|
| IP | 192.168.3.250/22 | 192.168.3.252/24 | 192.168.3.253/24 |
| isc-dhcp-server | 4.4.3, **active** | NOT installed | NOT installed |
| Webmin | 2.652 (port 10000, SSL) | NOT installed | NOT installed |
| DHCP leases | **532 active** | — | — |
| DHCP config | `/etc/dhcp/dhcpd.conf` | — | — |
**Source DHCP config summary:**
- Subnet: `192.168.0.0/22` (covers .0.x/.1.x/.2.x/.3.x)
- Pool range: `192.168.0.1 — 192.168.3.200`
- Router/GW: `192.168.3.254`
- DNS: `192.168.3.250` (netboot itself — **changing to 252/253**)
- Domain: `knel.net`
- Lease times: 600s default, 7200s max
- DDNS: `ddns-update-style none` (globally disabled)
- ~37 host reservations (MAC → fixed IP)
- 3 global-scope host declarations (outside subnet)
## 2. What changes in the migrated config
| Setting | Old (netboot) | New (netinfra) | Why |
|---------|--------------|----------------|-----|
| DNS servers handed out | `192.168.3.250` | `192.168.3.252, 192.168.3.253` | Point clients at the netinfra Pi-hole pair (netboot will be decommissioned) |
| Failover | none (single server) | ISC DHCP failover-peer | Primary/secondary HA |
| DHCP interface | `eth0` | `ens18` | Different NIC name on netinfra VMs |
| Netmask on servers | /22 | **fix to /22** (currently /24) | Match the /22 network |
Everything else (subnet, pool range, reservations, lease times, domain) stays
identical. The host reservation list is copied verbatim.
## 3. Configs prepared
| File | Destination |
|------|-------------|
| [`dhcp/dhcpd-primary.conf`](dhcp/dhcpd-primary.conf) | `/etc/dhcp/dhcpd.conf` on netinfra-01 |
| [`dhcp/dhcpd-secondary.conf`](dhcp/dhcpd-secondary.conf) | `/etc/dhcp/dhcpd.conf` on netinfra-02 |
Both are identical except for the `failover peer` block (primary vs secondary).
## 4. Execution sequence
### Phase 1 — pre-flight (non-disruptive)
1. **Snapshot sectestbed-sandbox** (VMID 6000 on tsys5) — rollback point
2. **Fix netmask** on netinfra-01/02: `/24``/22` in `/etc/network/interfaces`
(then `ifdown ens18 && ifup ens18`, or `systemctl restart networking`)
3. Verify Tailscale still reachable after netmask change
### Phase 2 — install prerequisites (non-disruptive, nothing serving yet)
On **both** netinfra-01 and netinfra-02:
1. `apt-get update && apt-get install -y isc-dhcp-server`
2. Install Webmin via official setup script
3. Copy `dhcpd-{primary,secondary}.conf``/etc/dhcp/dhcpd.conf`
4. Write `/etc/default/isc-dhcp-server` with `INTERFACESv4="ens18"`
5. Syntax check: `dhcpd -t -cf /etc/dhcp/dhcpd.conf`
6. **Do NOT start the DHCP service yet**
### Phase 3 — cutover (the critical moment)
**Goal:** minimize the DHCP gap. Existing clients hold valid leases for 600
seconds — they won't notice a brief outage.
1. **Stop netboot DHCP:** `ssh netboot 'sudo systemctl stop isc-dhcp-server'`
2. **Immediately start netinfra-01:** `systemctl start isc-dhcp-server`
3. **Immediately start netinfra-02:** `systemctl start isc-dhcp-server`
4. **Verify failover partnership:** check `journalctl -u isc-dhcp-server` on both
— look for "failover peer pfv-dhcp: I move from recover to recover-done"
5. Wait ~10 seconds for lease sync
### Phase 4 — test
1. On sectestbed-sandbox: edit `/etc/network/interfaces` to DHCP
2. `dhclient -r` (release) then `dhclient` (request new lease)
3. Verify IP assigned, correct DNS, correct gateway
4. If fails: `qm rollback 6000 <snapshot>` and debug
5. If succeeds: change back to static (sandbox's normal config)
### Phase 5 — cleanup
1. **Disable netboot DHCP:** `ssh netboot 'sudo systemctl disable isc-dhcp-server'`
2. Verify Webmin DHCP module visible at `https://192.168.3.252:10000`
3. Document + commit
## 5. Failover mechanics (reference)
ISC DHCP failover uses a TCP connection on port **647** between the two servers.
- **Primary** owns the pool initially and can assign new leases immediately.
- **Secondary** enters `recover` state, syncs the lease database from primary,
then transitions to `recover-done``normal`.
- If primary dies, secondary continues serving after the MCLT timeout (600s).
- If secondary dies, primary continues serving alone.
- **split 128** = load balance 50/50 (primary handles even hashes, secondary odd).
- Host reservations (fixed-address) are NOT subject to failover — both servers
can independently assign reserved IPs.
> **Documentation moved to Discourse — the canonical source of truth.**
>
> **DHCP migration to netinfra-01/02**
>
> **Read it here:** https://community.turnsys.com/t/306
>
> *Migrated 2026-08-06. This file is kept as a pointer for git-browsing context.
> Do not update content here — edit the Discourse wiki topic instead.*