# DHCP Migration: pfv-netboot → pfv-netinfra-01/02 **Date:** 2026-07-29 **Status:** ✅ **COMPLETE** — both servers in `normal` failover state, 333 active leases load-balanced. Netboot DHCP stopped + disabled. > **Safety nets:** pfv-netboot, pfv-netinfra-01/02, sectestbed-sandbox, and > ultix-streaming all have static IPs (reachable via Tailscale even if DHCP > fails). sectestbed-sandbox snapshot `pre-dhcp-migration` exists as rollback. ## Migration results (2026-07-29 17:20 CDT) | Check | Result | |-------|--------| | netinfra-01 (primary) | `active`, failover `normal`, 244 active leases | | netinfra-02 (secondary) | `active`, failover `normal`, 89 active leases | | Failover partnership | "Both servers normal" — load-balanced at split=128 | | DHCP test (sandbox) | Lease obtained from 192.168.3.252, IP 192.168.3.50/22 | | DNS handed out | `192.168.3.252, 192.168.3.253` (netinfra pair) | | NTP handed out | `192.168.3.252, 192.168.3.253` (netinfra pair) | | Domain | `knel.net` | | Gateway | `192.168.3.254` | | Webmin | Port 10000 (SSL) on both nodes; DHCP module configured | | Netboot DHCP | Stopped + disabled (won't start on reboot) | | Netmask fix | Both nodes corrected /24 → /22 | ## 1. Current state (discovered 2026-07-29) | | pfv-netboot (SOURCE) | pfv-netinfra-01 (TARGET primary) | pfv-netinfra-02 (TARGET secondary) | |---|---|---|---| | IP | 192.168.3.250/22 | 192.168.3.252/24 | 192.168.3.253/24 | | isc-dhcp-server | 4.4.3, **active** | NOT installed | NOT installed | | Webmin | 2.652 (port 10000, SSL) | NOT installed | NOT installed | | DHCP leases | **532 active** | — | — | | DHCP config | `/etc/dhcp/dhcpd.conf` | — | — | **Source DHCP config summary:** - Subnet: `192.168.0.0/22` (covers .0.x/.1.x/.2.x/.3.x) - Pool range: `192.168.0.1 — 192.168.3.200` - Router/GW: `192.168.3.254` - DNS: `192.168.3.250` (netboot itself — **changing to 252/253**) - Domain: `knel.net` - Lease times: 600s default, 7200s max - DDNS: `ddns-update-style none` (globally disabled) - ~37 host reservations (MAC → fixed IP) - 3 global-scope host declarations (outside subnet) ## 2. What changes in the migrated config | Setting | Old (netboot) | New (netinfra) | Why | |---------|--------------|----------------|-----| | DNS servers handed out | `192.168.3.250` | `192.168.3.252, 192.168.3.253` | Point clients at the netinfra Pi-hole pair (netboot will be decommissioned) | | Failover | none (single server) | ISC DHCP failover-peer | Primary/secondary HA | | DHCP interface | `eth0` | `ens18` | Different NIC name on netinfra VMs | | Netmask on servers | /22 | **fix to /22** (currently /24) | Match the /22 network | Everything else (subnet, pool range, reservations, lease times, domain) stays identical. The host reservation list is copied verbatim. ## 3. Configs prepared | File | Destination | |------|-------------| | [`dhcp/dhcpd-primary.conf`](dhcp/dhcpd-primary.conf) | `/etc/dhcp/dhcpd.conf` on netinfra-01 | | [`dhcp/dhcpd-secondary.conf`](dhcp/dhcpd-secondary.conf) | `/etc/dhcp/dhcpd.conf` on netinfra-02 | Both are identical except for the `failover peer` block (primary vs secondary). ## 4. Execution sequence ### Phase 1 — pre-flight (non-disruptive) 1. **Snapshot sectestbed-sandbox** (VMID 6000 on tsys5) — rollback point 2. **Fix netmask** on netinfra-01/02: `/24` → `/22` in `/etc/network/interfaces` (then `ifdown ens18 && ifup ens18`, or `systemctl restart networking`) 3. Verify Tailscale still reachable after netmask change ### Phase 2 — install prerequisites (non-disruptive, nothing serving yet) On **both** netinfra-01 and netinfra-02: 1. `apt-get update && apt-get install -y isc-dhcp-server` 2. Install Webmin via official setup script 3. Copy `dhcpd-{primary,secondary}.conf` → `/etc/dhcp/dhcpd.conf` 4. Write `/etc/default/isc-dhcp-server` with `INTERFACESv4="ens18"` 5. Syntax check: `dhcpd -t -cf /etc/dhcp/dhcpd.conf` 6. **Do NOT start the DHCP service yet** ### Phase 3 — cutover (the critical moment) **Goal:** minimize the DHCP gap. Existing clients hold valid leases for 600 seconds — they won't notice a brief outage. 1. **Stop netboot DHCP:** `ssh netboot 'sudo systemctl stop isc-dhcp-server'` 2. **Immediately start netinfra-01:** `systemctl start isc-dhcp-server` 3. **Immediately start netinfra-02:** `systemctl start isc-dhcp-server` 4. **Verify failover partnership:** check `journalctl -u isc-dhcp-server` on both — look for "failover peer pfv-dhcp: I move from recover to recover-done" 5. Wait ~10 seconds for lease sync ### Phase 4 — test 1. On sectestbed-sandbox: edit `/etc/network/interfaces` to DHCP 2. `dhclient -r` (release) then `dhclient` (request new lease) 3. Verify IP assigned, correct DNS, correct gateway 4. If fails: `qm rollback 6000 ` and debug 5. If succeeds: change back to static (sandbox's normal config) ### Phase 5 — cleanup 1. **Disable netboot DHCP:** `ssh netboot 'sudo systemctl disable isc-dhcp-server'` 2. Verify Webmin DHCP module visible at `https://192.168.3.252:10000` 3. Document + commit ## 5. Failover mechanics (reference) ISC DHCP failover uses a TCP connection on port **647** between the two servers. - **Primary** owns the pool initially and can assign new leases immediately. - **Secondary** enters `recover` state, syncs the lease database from primary, then transitions to `recover-done` → `normal`. - If primary dies, secondary continues serving after the MCLT timeout (600s). - If secondary dies, primary continues serving alone. - **split 128** = load balance 50/50 (primary handles even hashes, secondary odd). - Host reservations (fixed-address) are NOT subject to failover — both servers can independently assign reserved IPs.