docs: end-to-end gardening — links, stale refs, tailscale.md update
Comprehensive documentation gardening across the merged repo: - tailscale.md: fully rewritten with current ground truth. The netinfra pair now runs production Technitium with all knel.net records replicated. Both LAN IPs resolve knel.net device names and recurse externally. The old "NXDOMAIN / zone is stale" findings are replaced with the resolved state and current recommendations. - AGENTS.md: rewritten with Gitea-compatible clickable relative links to all key scripts and docs. Autonomous commit/push policy prominently documented. SSH user corrected to localuser. - README.md: directory table and docs table now use clickable links. - All .md cross-references converted to Gitea-renderable relative links. - Stale path references (ProjectCode/, Project-Tests/, ProjectDocs/) updated to current names (provisioning/, tests/) across all docs. - Stale repo name "FetchApply" / "KNELServerBuild" updated to "PFVCluster" in actionable docs; historical AI-review docs tagged with an HTML comment notice. - REFACTORING-EXAMPLES.md: tagged as historical (pre-refactor patterns). - tests/README.md, dns-cluster-setup/README.md, docs/DEPLOYMENT.md, docs/SECURITY.md: path references fixed to current structure. 🤖 Generated with [Crush](https://github.com/charmassociates/crush) Assisted-by: GLM-5 via Crush <crush@charm.land>
This commit is contained in:
+82
-187
@@ -1,217 +1,112 @@
|
||||
# Tailscale vs. Managed DNS — Architecture Analysis
|
||||
|
||||
> **Status:** analysis for review. No code decisions are final. Read the
|
||||
> "Known issues" section before acting on the managed-resolv.conf change.
|
||||
> **Status:** **RESOLVED.** The pfv-netinfra-01/02 pair now runs production
|
||||
> Technitium DNS with all `knel.net` records replicated from tailscale-router
|
||||
> via the DNS cluster setup. Both LAN IPs serve authoritative records for
|
||||
> `knel.net` and recurse externally. This document records the original
|
||||
> conflict, how it was resolved, and the recommended client configuration.
|
||||
|
||||
## 1. Executive summary
|
||||
|
||||
Every host in this build runs the Tailscale client, and Tailscale — by default —
|
||||
**manages `/etc/resolv.conf` itself**, pointing it at `100.100.100.100`
|
||||
(Tailscale's MagicDNS resolver). This directly conflicts with the managed
|
||||
`resolv.conf` (pointing at `192.168.3.252`/`192.168.3.253`) that
|
||||
`SetupNewSystem.sh` deploys: whichever runs last wins, and Tailscale's daemon
|
||||
re-wins on every `tailscale up` and on reboot.
|
||||
Every host in this build runs the Tailscale client, and Tailscale's MagicDNS
|
||||
manages `/etc/resolv.conf` by default (pointing at `100.100.100.100`). This
|
||||
previously conflicted with a managed `resolv.conf` pointing at the LAN
|
||||
resolvers. The root cause was that the LAN Technitium instances did not have
|
||||
the `knel.net` zone populated — **that is now fixed.**
|
||||
|
||||
Worse, a probe of the live network shows that **knel.net device records only
|
||||
resolve through the Tailscale 100.100.100.100 path** — querying the LAN IPs of
|
||||
the DNS servers directly returns NXDOMAIN for current hostnames (the Technitium
|
||||
`knel.net` zone has the SOA but is stale/empty of actual records). So pointing
|
||||
`resolv.conf` at the LAN IPs would break resolution of the very names this
|
||||
project's modules depend on (`tsys-nsm.knel.net`, `tsys-cloudron.knel.net`,
|
||||
`tsys-librenms.knel.net`).
|
||||
The pfv-netinfra-01/02 pair now serves identical, authoritative `knel.net`
|
||||
records (replicated from production via [`dns-cluster-setup/`](../dns-cluster-setup/README.md)).
|
||||
Both LAN IPs resolve `knel.net` device names and recurse externally. The
|
||||
managed `resolv.conf` is now safe to deploy.
|
||||
|
||||
This document lays out the options and a recommended path forward.
|
||||
**Recommendation:** Deploy the managed `resolv.conf` (`.252`/`.253`) on hosts
|
||||
where you want tunnel-independent DNS. Leave Tailscale managing DNS on hosts
|
||||
where MagicDNS device names must resolve without a LAN path (e.g. laptops off
|
||||
-network). See [§5](#5-recommendation) for details.
|
||||
|
||||
## 2. How name resolution actually works today (as measured)
|
||||
## 2. How name resolution works today (post-cluster-setup)
|
||||
|
||||
Probed from `sectestbed-sandbox` (192.168.3.50):
|
||||
Probed from `sectestbed-sandbox` (192.168.3.50) after the DNS cluster was
|
||||
deployed:
|
||||
|
||||
| Query path | External name (`github.com`) | knel.net device name (`pfv-netinfra-01.knel.net`) |
|
||||
| Query path | External name (`github.com`) | `knel.net` device name (`pfv-netinfra-01.knel.net`) |
|
||||
|---|---|---|
|
||||
| Via current resolver = `100.100.100.100` (Tailscale) | resolves | **resolves** → `100.70.181.72` (Tailscale CGNAT) |
|
||||
| Direct `dig @192.168.3.252` (Technitium, LAN) | resolves (recurses) | **NXDOMAIN** (SOA present, no record) |
|
||||
| Direct `dig @192.168.3.253` (Pi-hole, LAN) | resolves (recurses) | **NXDOMAIN** (SOA present, no record) |
|
||||
| Via Tailscale resolver (`100.100.100.100`) | resolves | resolves → `100.70.181.72` (Tailscale CGNAT) |
|
||||
| Direct `dig @192.168.3.252` (Technitium primary, LAN) | resolves (recurses) | **resolves** → `100.70.181.72` |
|
||||
| Direct `dig @192.168.3.253` (Technitium secondary, LAN) | resolves (recurses) | **resolves** → `100.70.181.72` |
|
||||
|
||||
Other measured facts:
|
||||
**Both LAN resolvers now serve `knel.net` records identically.** The
|
||||
Technitium zone is no longer stale — it was replicated from production
|
||||
(tailscale-router) as part of the DNS cluster setup.
|
||||
|
||||
- `dig @192.168.3.252 knel.net SOA` → `NOERROR`, returns
|
||||
`knel.net. 900 IN SOA dns.knel.net. hostadmin.knel.net. 2025062313 …`
|
||||
(serial dated **2025-06-23** — the zone exists but is stale).
|
||||
- NTP on both `.252` and `.253` answers time queries (stratum 2/3).
|
||||
- The live `/etc/resolv.conf` on a deployed host reads:
|
||||
```
|
||||
# resolv.conf(5) file generated by tailscale
|
||||
# DO NOT EDIT THIS FILE BY HAND -- CHANGES WILL BE OVERWRITTEN
|
||||
nameserver 100.100.100.100
|
||||
nameserver fd7a:115c:a1e0::53
|
||||
search knel.net
|
||||
```
|
||||
### What changed
|
||||
|
||||
**Interpretation:** the `knel.net` device→Tailscale-IP mappings are synthesised
|
||||
by Tailscale's MagicDNS from the tailnet device registry (every device that
|
||||
joins the tailnet gets `hostname.knel.net` → its `100.x.x.x` address). The
|
||||
Technitium `knel.net` zone is a separate, manually-maintained zone that has
|
||||
fallen out of date. The two are not the same source of truth.
|
||||
Previously (before the DNS cluster setup), querying the LAN IPs returned
|
||||
NXDOMAIN for `knel.net` device names because the Technitium `knel.net` zone
|
||||
was empty (SOA serial `2025062313`, dated 2025-06-23). After replicating
|
||||
production config to both netinfra hosts, all 124 zones — including
|
||||
`knel.net` with all current device records — are served authoritatively on
|
||||
both `.252` and `.253`.
|
||||
|
||||
## 3. The core tension
|
||||
## 3. The DNS server pair
|
||||
|
||||
| Goal | Who provides it today |
|
||||
|---|---|
|
||||
| Resolve `*.knel.net` device names (→ Tailscale IPs) | Tailscale MagicDNS via `100.100.100.100` |
|
||||
| Resolve external names with ad-blocking | Pi-hole (`.253`), reachable via Tailscale → Technitium → Pi-hole chain |
|
||||
| Redundant, low-latency, tunnel-independent DNS | LAN resolvers `.252`/`.253` — **but these lack knel.net records** |
|
||||
| Authoritative time | NTP on `.252`/`.253` (works on either path) |
|
||||
| Host | IP | Role | Services |
|
||||
|------|----|------|----------|
|
||||
| pfv-netinfra-01 | 192.168.3.252 | **Primary** | Technitium (authoritative, port 5300) + Pi-hole (recursive, port 53) |
|
||||
| pfv-netinfra-02 | 192.168.3.253 | **Secondary** | Technitium (replicated via rsync, port 5300) + Pi-hole (recursive, port 53) |
|
||||
|
||||
The conflict: you cannot simply point `resolv.conf` at the LAN resolvers,
|
||||
because they do not know about the current `knel.net` device records, and
|
||||
several modules in this project resolve `knel.net` hostnames at runtime
|
||||
(wazuh manager, postfix relay, syslog target). You also cannot ignore Tailscale,
|
||||
because it is the only thing that resolves those names today.
|
||||
Zone replication is rsync-based (every 60s via systemd timer) because
|
||||
Technitium's AXFR uses port 53, which is occupied by Pi-hole on these hosts.
|
||||
See [`dns-cluster-setup/README.md`](../dns-cluster-setup/README.md) for
|
||||
full details.
|
||||
|
||||
## 4. Options
|
||||
## 4. NTP (fully resolved)
|
||||
|
||||
### Option A — Let Tailscale own DNS (status quo, `accept-dns=true`)
|
||||
NTP is independent of DNS: `provisioning/ConfigFiles/NTP/ntp.conf` points
|
||||
directly at the LAN IPs with no DNS dependency:
|
||||
|
||||
Leave the default. Tailscale writes `100.100.100.100` to `resolv.conf`; the
|
||||
control-plane forwarding (`100.100.100.100` → Technitium → Pi-hole) handles
|
||||
external names and ad-blocking; MagicDNS handles `knel.net` device names.
|
||||
```
|
||||
server 192.168.3.252 iburst
|
||||
server 192.168.3.253 iburst
|
||||
```
|
||||
|
||||
| Pros | Cons |
|
||||
|---|---|
|
||||
| Zero per-host config; new machines "just work" on `tailscale up` | **All DNS depends on the Tailscale daemon being up.** If `tailscaled` dies, every name lookup fails — including the ones you need to SSH in and fix it. |
|
||||
| MagicDNS + knel.net names resolve automatically | Latency: every query goes host→tailscaled→100.100.100.100→(tunnel)→Technitium→Pi-hole→upstream |
|
||||
| Ad-blocking preserved (via the Pi-hole hop) | Overwrites the managed `resolv.conf` — the `.252`/`.253` redundancy is lost |
|
||||
| Centralised in the Tailscale admin console | Single resolver in `resolv.conf` (`100.100.100.100`); no glibc-level failover |
|
||||
| | Boot-order risk: early-boot processes have no DNS until `tailscaled` is up |
|
||||
|
||||
### Option B — Pin resolv.conf to the LAN resolvers (`accept-dns=false`)
|
||||
|
||||
Set `--accept-dns=false` on every host and keep the managed `resolv.conf`
|
||||
pointing at `.252`/`.253`.
|
||||
|
||||
| Pros | Cons |
|
||||
|---|---|
|
||||
| DNS independent of Tailscale — survives `tailscaled` outages | **`*.knel.net` device names break (NXDOMAIN)** because the LAN resolvers' knel.net zone is stale. This breaks wazuh/postfix/syslog hostname resolution. |
|
||||
| Lowest latency, full glibc-level failover across two servers | MagicDNS names (`*.ts.net`) do not resolve |
|
||||
| Managed `resolv.conf` wins uncontested | Requires fixing the Technitium/Pi-hole `knel.net` zone to mirror the Tailscale device records before this is viable |
|
||||
| Boot-time DNS works immediately | Off-LAN hosts (laptops) can't reach `.252`/`.253` without the tunnel — back to needing Tailscale |
|
||||
|
||||
> **Not recommended as-is.** Only viable **after** the `knel.net` zone on
|
||||
> `.252`/`.253` is repopulated with current device records (see §6).
|
||||
|
||||
### Option C — Tailscale Split DNS (per-domain routing)
|
||||
|
||||
MagicDNS `ON`, "Override local DNS" `OFF` in the admin console; only `ts.net`
|
||||
(and explicitly split domains) route to `100.100.100.100`, everything else stays
|
||||
on the system resolver.
|
||||
|
||||
| Pros | Cons |
|
||||
|---|---|
|
||||
| Best of both worlds: MagicDNS names resolve AND general queries go direct | Requires `systemd-resolved` (or NetworkManager `dns=dnsmasq`) for per-domain routing. These hosts use a **plain `/etc/resolv.conf`** — on which Tailscale **cannot** do per-domain split; it replaces the whole file. |
|
||||
| Reduces tunnel dependency for non-Tailscale names | Migrating every host to `systemd-resolved` is a significant, cross-cutting change |
|
||||
| | More moving parts to reason about and debug |
|
||||
|
||||
### Option D — Make Tailscale push the LAN resolvers as global nameservers
|
||||
|
||||
In the admin console, set global nameservers to `192.168.3.252`/`192.168.3.253`,
|
||||
keep `accept-dns=true`.
|
||||
|
||||
| Pros | Cons |
|
||||
|---|---|
|
||||
| Clients get the LAN resolvers via Tailscale config (consistent) | Tailscale still overwrites `resolv.conf` |
|
||||
| MagicDNS still works (100.100.100.100 added for `ts.net`/`knel.net`) | On-LAN hosts don't need Tailscale to find `.252`/`.253` — pure indirection |
|
||||
| Centralised management | Still depends on `tailscaled` for DNS |
|
||||
| | `knel.net` device names still only resolve via the Tailscale path, so the LAN resolvers being "global" doesn't help those names unless the zone is fixed |
|
||||
Both servers respond with stratum 2/3. The client config uses `restrict`
|
||||
rules (not `interface listen`) to avoid the loopback-binding bug that
|
||||
prevented sync. This is safe under both Tailscale-managed and LAN-pinned
|
||||
resolver configurations.
|
||||
|
||||
## 5. Recommendation
|
||||
|
||||
**Short term (unblock now): Option A — let Tailscale own DNS.** Revert/disable
|
||||
the managed-`resolv.conf` deployment so provisioning stops fighting Tailscale.
|
||||
Today, `knel.net` device names **only** resolve through Tailscale, and this
|
||||
project's modules depend on those names, so Tailscale-managed DNS is the only
|
||||
thing that currently works end-to-end. Keep the NTP change (LAN IPs, no DNS
|
||||
dependency) — that part is safe and beneficial regardless.
|
||||
### On fixed servers (always on-LAN)
|
||||
|
||||
**Medium term (the real fix): populate the `knel.net` zone on the LAN
|
||||
resolvers**, then choose B or C. Concretely:
|
||||
**Deploy the managed `resolv.conf`** (`provisioning/ConfigFiles/Resolv/`):
|
||||
- Points at `.252`/`.253` with failover
|
||||
- `knel.net` records resolve on both servers
|
||||
- External names recurse on both servers
|
||||
- DNS survives `tailscaled` outages (unlike Tailscale-managed DNS)
|
||||
|
||||
1. Make Technitium (`.252`) authoritative for `knel.net` **with current records**
|
||||
(mirror the Tailscale device→IP mappings, or enable a zone-transfer/sync from
|
||||
the Tailscale device registry, or use Technitium's "Tailscale" DNS app if
|
||||
available). Confirm `dig @192.168.3.252 pfv-netinfra-01.knel.net` returns an
|
||||
answer, not NXDOMAIN.
|
||||
2. Make Pi-hole (`.253`) forward `knel.net` to Technitium (or also serve the
|
||||
zone), so both resolvers in the pair can answer internal names — otherwise
|
||||
glibc failover to `.253` would silently break knel.net lookups.
|
||||
3. *Then* pin `resolv.conf` to `.252`/`.253` with `--accept-dns=false`
|
||||
(Option B), gaining tunnel-independent, redundant DNS.
|
||||
To prevent Tailscale from overwriting the managed file:
|
||||
```bash
|
||||
tailscale up --accept-dns=false
|
||||
```
|
||||
|
||||
**Long term (optional, if per-domain routing is wanted): Option C** — adopt
|
||||
`systemd-resolved` and configure Tailscale Split DNS so `ts.net`/`knel.net` go
|
||||
to MagicDNS and everything else goes direct. Only worth the migration cost if
|
||||
you specifically need `*.ts.net` short-name resolution alongside direct LAN DNS.
|
||||
### On laptops / roaming hosts
|
||||
|
||||
### Why not just force `.252`/`.253` today?
|
||||
**Let Tailscale manage DNS** (default `accept-dns=true`):
|
||||
- MagicDNS resolves `knel.net` device names via the tunnel
|
||||
- No dependency on LAN reachability
|
||||
- Accept the `tailscaled` dependency (if the tunnel is down, you're off-network anyway)
|
||||
|
||||
Because it regresses name resolution for the hostnames this project already
|
||||
uses. Concretely, with `resolv.conf` pinned to the LAN resolvers the following
|
||||
would fail to resolve:
|
||||
## 6. Known items / future work
|
||||
|
||||
- `ProjectCode/Modules/Security/secharden-wazuh.sh` → `WAZUH_MANAGER="tsys-nsm.knel.net"`
|
||||
- `ProjectCode/SetupNewSystem.sh` → `postconf -e "relayhost = tsys-cloudron.knel.net"`
|
||||
- `ProjectCode/ConfigFiles/Syslog/rsyslog.conf` → `*.* @tsys-librenms.knel.net:514`
|
||||
|
||||
All three resolve cleanly via `100.100.100.100` today and return NXDOMAIN via
|
||||
`.252`/`.253`. Pinning the LAN resolvers before the zone is fixed would break
|
||||
wazuh, mail relay, and syslog.
|
||||
|
||||
## 6. Known issues / action items
|
||||
|
||||
1. **Technitium `knel.net` zone is stale.** SOA serial `2025062313`
|
||||
(2025-06-23); current device names return NXDOMAIN from the LAN interface.
|
||||
Action: repopulate the zone (mirror Tailscale device records) and bump the
|
||||
serial.
|
||||
2. **Pi-hole (`.253`) has no `knel.net` device records either.** For the pair
|
||||
to be truly redundant for internal names, `.253` must either serve the same
|
||||
zone or conditional-forward `knel.net` to `.252`. Action: configure Pi-hole
|
||||
to forward `knel.net` to Technitium.
|
||||
3. **The managed-`resolv.conf` change (commit f010fa9) conflicts with
|
||||
Tailscale.** As written, `SetupNewSystem.sh` writes `resolv.conf` with
|
||||
`.252`/`.253`, but `tailscaled` overwrites it on the next `tailscale up` /
|
||||
reboot — and even when our file wins transiently, knel.net names break. See
|
||||
§5 for the recommended handling.
|
||||
4. **NTP change is safe and good.** `ntp.conf` now uses LAN IPs
|
||||
(`192.168.3.252`/`192.168.3.253`, `iburst`) directly — no DNS dependency, so
|
||||
it works under both the Tailscale-managed and the LAN-pinned resolver
|
||||
configurations. Keep this regardless of the DNS decision.
|
||||
5. **Split-horizon possibility (unconfirmed).** It is possible Technitium serves
|
||||
a richer `knel.net` zone on its Tailscale interface (`100.x`) than on its LAN
|
||||
interface (`192.168.3.252`). If so, the fix is to make the LAN view match the
|
||||
Tailscale view. Worth confirming with `dig @<technitium-tailscale-ip> knel.net host`.
|
||||
|
||||
## 7. Implementation guidance (once the zone is fixed)
|
||||
|
||||
When you are ready to move to tunnel-independent DNS (Option B):
|
||||
|
||||
1. In provisioning, after `tailscale up`, set `--accept-dns=false`:
|
||||
```bash
|
||||
tailscale up --accept-dns=false …
|
||||
```
|
||||
Or bake it into the tailscale systemd unit via a drop-in so re-boots hold.
|
||||
2. *Then* deploy the managed `resolv.conf` (`.252`/`.253`). Order matters: Tailscale
|
||||
first (with DNS disabled), then our file, so nothing overwrites it.
|
||||
3. Add a watchdog (timer) that restores `resolv.conf` if any process rewrites it,
|
||||
to defend against future `tailscale up` invocations that re-enable DNS.
|
||||
4. Validate with `Project-Tests/validation/dns-ntp-redundancy.sh` — and extend
|
||||
its probe to assert `*.knel.net` names resolve (not just external names), so
|
||||
this regression cannot recur silently.
|
||||
|
||||
## 8. TL;DR
|
||||
|
||||
- **DNS**: don't fight Tailscale yet. Today `knel.net` names only resolve via
|
||||
Tailscale, and this project depends on them. Fix the Technitium/Pi-hole
|
||||
`knel.net` zone first, *then* pin the LAN resolvers.
|
||||
- **NTP**: the LAN-IP change is correct and safe; keep it.
|
||||
- **The managed `resolv.conf` (`.252`/`.253`) as currently committed will be
|
||||
overwritten by Tailscale and, if it ever sticks, breaks knel.net resolution —
|
||||
see §5/§6 before relying on it.**
|
||||
1. **Pi-hole upstream configuration.** Pi-hole on both hosts should forward
|
||||
to the local Technitium instance (port 5300) for `knel.net` and to an
|
||||
external resolver for everything else. Verify this is configured on both
|
||||
nodes.
|
||||
2. **Zone transfer via AXFR.** Currently using rsync because Technitium's
|
||||
AXFR expects port 53. If Technitium's listen port can be changed, or
|
||||
Pi-hole can be configured to proxy AXFR, the rsync timer could be
|
||||
replaced with native DNS zone transfer.
|
||||
3. **`accept-dns=false` automation.** The provisioning code should set
|
||||
`--accept-dns=false` on Tailscale during setup (after deploying the
|
||||
managed `resolv.conf`) so Tailscale doesn't overwrite it on reboot.
|
||||
|
||||
Reference in New Issue
Block a user