340 Commits
Author SHA1 Message Date
reachableceo 4eae30c979 Merge pull request 'retire .creds references in AGENTS.md (vault via sm)' (#1) from ic-builder/creds-retirement into main 2026-09-06 22:33:59 +00:00
ic-builder ceb5c3fe91 docs: retire .creds pattern references (vault via sm; identities rule)
Ticket: https://projects.knownelement.com/issues/829
2026-09-06 17:32:06 -05:00
mrcharles a73591045e [#769] framework: set-e-safe cn_count fallback
https://projects.knownelement.com/issues/769
2026-09-05 07:36:47 -05:00
mrcharles 37b54db824 [#769] framework fix: cn_count double-count in compose check
https://projects.knownelement.com/issues/769
2026-09-05 07:17:22 -05:00
mrcharles 2161ce64c2 [#769] retire questions files — open items ticketed in Redmine (NeedsInput)
v10/v11 archived under archive/questions/; carried opens mapped to
tickets #800-#807.
https://projects.knownelement.com/issues/769
2026-09-05 05:28:25 -05:00
mrcharles 080f5c9291 [#769] stub k8s/awx — split to KNEL/k8s + KNEL/awx
https://projects.knownelement.com/issues/769
2026-09-05 05:20:08 -05:00
reachableceo 9649b43b67 approved 2026-09-05 01:34:15 -05:00
mrcharles 4d7287c005 docs(questions): v11 — ACME DNS-01 TSIG + step-ca go; legacy-* exempt [#697][#771]
https://projects.knownelement.com/issues/697
2026-09-04 17:31:52 -05:00
mrcharles 8d28d1c064 feat(sweep): backup retention — keep newest 2 automatic backups [#777] 2026-09-04 16:35:49 -05:00
mrcharles 18c590ba2a feat(k8s): nvidia device-plugin manifest SoR — GPU verified E2E [#763][#769]
k8s now in the O&M lane (Charles 09-04). Verification: capacity/
allocatable nvidia.com/gpu=1 on wnode-tsys3; CUDA pod scheduled and
nvidia-smi clean in-container. Q2 sweep: zero node_exporters on all
PVE hosts (Beszel owns host metrics).

https://projects.knownelement.com/issues/763
2026-09-04 15:11:00 -05:00
mrcharles 61cdc8b661 feat(sweep): daily maintenance sweep script [#777]
HA health, sensor freshness (watchdog classes), deploy pipeline,
runner, NUT both UPSes, DNS, disk. Exit = failure count.
2026-09-04 14:21:41 -05:00
mrcharles ad69416a19 docs: runner org-scoped to KNEL + ultix label [#784] 2026-09-04 13:14:50 -05:00
mrcharles abd677a32e docs: runner label ultix — workflows must run on ultix [#784] 2026-09-04 12:27:51 -05:00
mrcharles c89ca6c47a docs(hooks): CHANGE GATE message points at mglpi [#767]
CR filing one-liner now the Go connector (mopac-glpi-go) instead of the
deprecated bash shim. Functional gate re-verified after edit.

https://projects.knownelement.com/issues/767#note-4191
2026-09-04 09:01:04 -05:00
mrcharles f6cd430699 docs: OQ8 answered — pfv-bms to KNELBMS rename approved [#785]
Founder 2026-09-04: KNELBMS becomes the fleet gold standard for
SDLC/CI-CD/prod stability. Staged plan in #785.
2026-09-04 07:54:49 -05:00
mrcharles a3bd8bb7d4 docs: Key Scripts table — pve-snapshot, garden, audit-fleet, kuma suite [#784] 2026-09-04 07:48:15 -05:00
mrcharles d258fc2bb5 docs: CI/CD fleet standard — gitea runner + repo workflow duty [#784]
https://community.turnsys.com/t/333
2026-09-04 07:26:54 -05:00
mrcharles a024b78cde feat(hooks): CHANGE GATE — prod remote ops require .crush/active-cr [#767]
Q3 ruling 2026-09-04: Tier 2 of the t/325 ladder is mechanical. Prod-
target remote.sh/remote-dns.sh invocations block without a filed GLPI
Change id; dev lanes (sectestbed/preprod/test/tsys5/sandbox/kali) exempt;
CR tooling exempt (chicken-and-egg). 7-case matrix green, shellcheck clean.

https://projects.knownelement.com/issues/767#note-4191
2026-09-04 07:16:24 -05:00
mrcharles cb4ef7e8e3 docs: field lesson — bridge re-enslave + wedged nconnect NFS recovery [#780]
https://projects.knownelement.com/issues/769#note-4152
2026-09-04 06:45:35 -05:00
mrcharles 66def2e311 feat(cm): auto-open Kuma maintenance windows on rollback (OQ5 YES) [#769]
kuma-maintenance.py (open/list/delete, single-strategy windows,
version-tolerant addMaintenance probing) + pve-snapshot.sh rollback
opens a 30m window before touching a VM when ~/.creds/uptime-kuma.env
exists. TZ pinned America/Chicago. Verified: open->list(active=True)->
delete live on the fleet Kuma.

https://projects.knownelement.com/issues/769#note-4152
2026-09-04 06:05:27 -05:00
mrcharles ce55409b9a docs: archive superseded question rounds to archive/questions [#769]
v10 is the current round. Old rounds preserved in git history and
archive/questions/.
2026-09-04 05:57:04 -05:00
mrcharles 823eb9b33d docs: AGENTS.md repoint to KNEL repos + stale-ref fixes [#769]
Completes dba54b2 (which staged deletions only — pathspec'd add missed
the content edits). Layout table, Key Scripts, DNS discipline, conman
section now point at KNEL/netinfra|facilities|perf|monitoring|inventory;
conman example switched from IP literal to pfv-tsys4 (DNS-names rule).

https://projects.knownelement.com/issues/769#note-4152
2026-09-04 05:52:06 -05:00
reachableceo c0f71756ee answered 2026-09-03 23:12:37 -05:00
mrcharles 6173be60b7 docs(questions): v10 — O&M split round + carried PMG items [#769]
New OQ1-OQ8 for the founder sitting; v9 PMG PQ5-PQ12 + v6 Q4 carried open.

https://projects.knownelement.com/issues/769#note-4152
2026-09-03 22:06:06 -05:00
mrcharles 8ba0ae1d6f feat(cm): pve-snapshot.sh — snapshot lifecycle for change management [#769]
create/list/rollback/clean with cm-<ticket>-<label>-<ts> naming,
protected-VM regex (pfv-bms/netinfra/PMG/PBS), Kuma-window gate on
rollback, retention cleaner, local audit log. Verified end-to-end on
tsys5 against throwaway VM 59999 (created, rolled back, destroyed).

https://projects.knownelement.com/issues/769#note-4152
2026-09-03 21:57:15 -05:00
mrcharles dba54b294f refactor: split domain tooling into dedicated KNEL repos [#769]
siem->KNEL/siem, netinfra->KNEL/netinfra, ca->KNEL/ca, oam->KNEL/monitoring,
cmdb->KNEL/inventory, dcinfra->KNEL/facilities, proxmox/perf->KNEL/perf;
unit tests moved with their code. AGENTS.md layout/paths repointed.
Full history retained here. Map: Redmine #769, Discourse t/331.

https://projects.knownelement.com/issues/769#note-4152
2026-09-03 21:49:50 -05:00
mrcharles ebd876b085 fix(rules): conflict-marker check false-positived on decorative banners [#769]
Tightened to exact git markers: 7-char = alone, <<<<<<< / >>>>>>> with ref.

https://projects.knownelement.com/issues/769#note-4152
2026-09-03 21:23:57 -05:00
mrcharles 041d311d1c docs(questions): C4 credentials + C5 scope + C6 docker rulings recorded [#767]
Answers landed 2026-09-03: Bitwarden/Vault = cred SoR (.creds bootstrap-only);
this chat = COO/C2-only (PM/policy/enablement), agent stack lives here;
docker(988) verified on all 7 TSG*/COS* officer accounts.
2026-09-03 17:27:54 -05:00
mrcharles 994959d4ad docs(questions): C4-C8 agent-stack + credentials-policy rulings [#767]
Detail: https://projects.knownelement.com/issues/767
2026-09-03 17:13:29 -05:00
mrcharles a2a45ea9d8 feat(bench): harness v1.1 — PSI latency, load, device identity context [#709]
Replaces unavailable latencytop; sysstat/iotop deployed fleet-wide
separately. Redmine: https://projects.knownelement.com/issues/709
2026-09-03 16:55:28 -05:00
mrcharles abb725be85 feat(oam): nginx vhost SoR — OAM docker UIs behind fleet TLS [#697]
4 subdomain vhosts proxying to local docker ports; SAN cert from fleet
CA; workstation chain-verified (ssl_verify=0 all four).

Detail: https://projects.knownelement.com/issues/697
2026-09-03 14:42:33 -05:00
mrcharles 8a4ee6f64f dns: oam UI subdomains — smokeping/netdisco/oxidized/unpoller → tsys-librenms [#697]
A + auto-PTR via dns-cli on primary; zone-snapshot + drift-check ALL
IN SYNC same session. Serves the fleet-TLS pass for the tsys-librenms
docker UIs.

Detail: https://projects.knownelement.com/issues/697
2026-09-03 14:39:24 -05:00
mrcharles ba4463c4d0 fix(oam): kuma inventory target-field bug + add-ping tooling [#343][#435]
kuma-inventory.py printed the junk url field (https://) instead of the
hostname for ping monitors — hid ~100 monitors from FQDN matching.
kuma-add-ping.py: idempotent ping-monitor adds (exists-check by name and
target). Matrix published to Discourse t/309 (canonical, #343).

Detail: https://projects.knownelement.com/issues/343
2026-09-03 14:34:44 -05:00
mrcharles a07d866116 fix(cmdb): canonical CI names — s7 rename map in seed generator [#705]
from-inventory.sh now applies the #307 s7 VM-name -> Tailscale-name
mapping (two-pass awk), so rr-middleware-01/preprod-proxmoxmailgw rows
emit canonical names. Seed regenerated; GLPI renames + dup purges done
in same session (78 CIs at parity).

Detail: https://projects.knownelement.com/issues/705
2026-09-03 14:22:05 -05:00
mrcharles 8666182d8b docs(cmdb): regenerate seed post-reconcile — 75 CIs, stale rows gone [#705]
Detail: https://projects.knownelement.com/issues/705
2026-09-03 14:05:04 -05:00
mrcharles 056412a311 feat(cmdb): inventory reconcile + GLPI software catalog seeder [#705]
#307 s6 rewritten post-verification: zone was ALREADY clean (zero DNS
deletes); 12 stale CIs purged from GLPI, 3 alive off-cluster systems
given proper CIs (tsys-cloudron, ultix-field/highside), 75-row seed
regenerated. software-catalog.sh: 21 products + 59 installs (wazuh
agents linked from live manager); version lookup fixed to client-side
parent filter (GLPI ignores softwares_id search param).

Detail: https://projects.knownelement.com/issues/705
2026-09-03 14:01:56 -05:00
mrcharles 4c461c14e1 refactor: move HA tooling + reference packs to KNEL/pfv-bms [#762]
HomeAssistant/ is now a redirect stub; AGENTS.md layout updated.
Zero HA impact — pfv-bms never loaded from PFVCluster.

Detail: https://projects.knownelement.com/issues/762
2026-09-03 13:34:55 -05:00
mrcharles 1620351028 feat(ca): artifact mirror live — pinned wazuh-agent debs on tsys-ca [#758]
dist/wazuh-agent/4.14.7-1/ (amd64+arm64+armhf) + SHA256SUMS, autoindex
enabled, workstation hash-verified against published manifest.
deploy-agent.sh now installs from the mirror (upstream fallback) and
auto-detects arch. #335: kali-tsys enrolled+active (30 total agents).

Detail: https://projects.knownelement.com/issues/758
2026-09-03 13:25:28 -05:00
mrcharles e27e3a06b8 feat(oam): codify garage-pdu-relay snmpd extend + timeout hardening [#733]
Deployed manually on tsys-librenms 2026-09-02 (founder relay ruling);
codified same-day rule. snmpget -t2 -r1 so a dark PDU answers in ~2s.
Deployed target: tsys-librenms:/usr/local/bin/garage-pdu-relay.

Detail: https://projects.knownelement.com/issues/733
2026-09-03 12:55:31 -05:00
mrcharles d3b9099d5f feat(mail): fleet mail-env audit + map; PMG VIP proposal [#696][#694]
Audit findings note (55/64 systems, read-only):
https://projects.knownelement.com/issues/696#note-4088

Map + deploy plan: https://community.turnsys.com/t/322
Key: relayhost empty fleet-wide (direct-to-MX today); prod PMG pair
already installed (VM 604/711, PMG 9.1, unclustered); VIP slot
proposed 192.168.3.249; open PQs in questions-v8.md.
2026-09-03 12:49:59 -05:00
mrcharles 7ca87c1200 docs(questions): C1-C3 CMDB change-control rulings needed [#705]
Detail: https://projects.knownelement.com/issues/705
2026-09-03 12:32:20 -05:00
mrcharles 412d706abd docs(questions): v7 — PMG design round PQ1-PQ7 [#696][#694]
Session-start note (ground truth + design questions):
https://projects.knownelement.com/issues/696#note-4086

No duplicate ticket: #696 (child of umbrella #694) already covers
PMG deploy + MX cut-over; session gated on #696.
2026-09-03 12:27:34 -05:00
mrcharles 07a0a3472b fix(cmdb): GLPI import fixes — input wrapper + Read-Only agent profile [#705]
Import executed live: 87 CIs seeded, scoped cmdb user (Read-Only)
verified read-ok / write-denied. Fixes: POST body needs {input:[...]}
(ERROR_BAD_ARRAY on bare array), profile swapped Technician ->
Read-Only per least-priv ruling, shape-tolerant user-exists check.

Detail: https://projects.knownelement.com/issues/705
2026-09-03 12:26:39 -05:00
mrcharles 379f7d376b docs(cmdb): regenerate seed — wnode-tsys5 ghost retired [#728]
#307 row annotated RETIRED (VM 500 gone from every node, not in the k8s
roster); seed regenerated via the converter. Also carries the ultix-offstage
retirement.
2026-09-03 11:46:36 -05:00
mrcharles b7eb9a5066 fix(bench): trixie libaio1 fallback in bench LXC installer [#709]
Redmine detail: https://projects.knownelement.com/issues/709
2026-09-03 11:43:48 -05:00
mrcharles d237e4a65d feat(monitoring): kuma notification tooling; app webhooks replace pushover [#435]
kuma-notifications.py (list/delete/usage). Deleted unused Pushover provider
'KNEL Alerts' per Charles — 211/211 monitors already on the ultix-sidecar +
Ultix-mini app webhooks. fleet-sync creation payload updated to the live
notification IDs.
2026-09-03 11:38:38 -05:00
mrcharles 4c2726b116 feat(cmdb): GLPI importer — agent user + seed import ready to fire [#705]
to-glpi.sh: initSession, creates scoped local 'cmdb' user (Technician @
root entity, random token -> ~/.creds/glpi-agent.env), then batch-imports
systems.csv as Computers (chunked, idempotent by name, --dry-run).
Fires the moment Charles drops the corrected GLPI_USER_TOKEN.
Pure bash/jq per house rules.
2026-09-03 09:59:32 -05:00
mrcharles b575a8e864 fix(perf): install libaio1t64 fallback for Debian 13 guests [#709] 2026-09-03 09:30:29 -05:00
mrcharles bf52f85ac5 docs(agents): field lessons from #684 PBS rollout + restore-drill incident
Meat: https://projects.knownelement.com/issues/684
2026-09-03 08:33:58 -05:00
mrcharles 7586605b70 fix(perf): bench jq direction bug + libaio + json note guard [#709]
- write tests read .write.* explicitly (jq // kept read=0 for writes)
- libaio1 installed everywhere (real aio engine, no sync cap)
- fio stdout notes stripped before jq parse

Ticket: https://projects.knownelement.com/issues/709
2026-09-03 07:53:37 -05:00
mrcharles 82fdb161cc feat(ca): fleet CA tooling — intermediate live on tsys-ca, first cert issued [#697]
ca-init/issue-cert/selftest (TDD loop, shellcheck clean); design doc on
Discourse t/320. Offline RSA-4096 root (Nitrokey ceremony later), 5y
intermediate, 825d SAN leaves. First cert: tsys-wazuh.knel.net (fingerprint
on the ticket note).
Meat: https://projects.knownelement.com/issues/697#note-4033
2026-09-03 07:52:07 -05:00
mrcharles 02bdeae9e7 feat(perf): standardized fleet benchmark artifact (fio/iperf3) [#709]
bench-run.sh: identical test shape everywhere (4k rand rw, 1M seq rw,
direct IO, grep-able BENCH| lines) - runs in guests, LXCs, containers.
install-bench.sh for Debian guests/LXCs; Dockerfile for registry image;
lxc-bench-setup.sh stands up the host-side bench LXC (proxmox = no docker).

Ticket: https://projects.knownelement.com/issues/709
2026-09-03 07:21:27 -05:00
mrcharles d75366edf5 feat(siem): agent fixes from fleet pass; IPv6 kill script [#335][#748]
deploy-agent.sh now self-heals the three postinst failure modes found on
PMG/PBS hosts (MANAGER_IP placeholder, missing wazuh user, root:root
ownership). scripts/disable-ipv6.sh: live sysctl + persisted conf, with
in-file rollback notes. Rollout state in #335/#748.
Meat: https://projects.knownelement.com/issues/335#note-4008
2026-09-03 07:13:28 -05:00
mrcharles 2485866ccd chore(dns): remove retired wnode-tsys5 ghost A record from knel.net [#737]
Node retired and replaced by ultix-streaming; tailnet-side entry still
requires admin-console removal (no TS API key in store). Zone snapshot
refreshed same-session per DNS sync rule.

Ticket: https://projects.knownelement.com/issues/737
2026-09-03 06:32:44 -05:00
mrcharles 06679d9a0c feat(siem): fleet rollout tooling — agents + PVE rsyslog forwarding [#335]
deploy-agent.sh (Ubuntu VMs) + deploy-rsyslog.sh (PVE hosts, TCP/514 via
Tailscale). Landed: agents on netinfra-01/02, librenms, awx (7 total
active on manager); rsyslog forwarding verified on all 7 PVE hosts
(persistent 514 sessions on the manager). Deferred: k8s nodes (k8s-chat
coordination), docker json-log caps (needs daemon restart window).
Meat + verification: https://projects.knownelement.com/issues/335#note-4006
2026-09-03 06:30:25 -05:00
mrcharles 5c5e173a1b feat(cmdb): seed tooling — inventory→CSV converter + 87-CI seed v0 [#705]
P0 of the CMDB/change-control plan (design: Discourse t/319). Regenerates
the GLPI seed dataset from the #307 inventory raw; re-runnable whenever
#307 changes. AGENTS.md: cmdb/ layout entry + Key Scripts row.

Meat + plan: https://projects.knownelement.com/issues/705#note-4004
2026-09-03 05:54:08 -05:00
mrcharles 6a38b4443c docs(agents): cross-linking house rule + mastodon post tool [#743][#441]
Clickable-refs mandate from Charles: ticket comments link commits,
commit bodies link the Redmine comment URL, Discourse links both —
one click between Redmine/Gitea/Discourse (GLPI/CMDB change control
will extend this). Mechanical body-URL check queued under #441.
scripts/mastodon/post.sh: zero-install curl poster; creds land in
~/.creds/mastodon.env (MASTODON_URL/MASTODON_TOKEN, write scope).
2026-09-03 05:28:57 -05:00
mrcharles 6c80b76b47 docs: questions v4 — pairing mechanism fork + gitea allowlist ask [#626][#734] 2026-09-02 23:10:07 -05:00
mrcharles 4a38e9bac8 docs(questions): v6 — Q5 resolved (separate roster repo, built); Q4 open (second human) [#345]
https://projects.knownelement.com/issues/345
2026-09-02 21:59:36 -05:00
mrcharles 7376763824 docs(questions): v5 — v4 answers recorded (relay hw, probe cleared, HA whitelist); PR approver questions [#345]
Founder answered Q1-Q3 2026-09-02; new Q4/Q5 ask for named PR
approvers and pfv-bms protection scope.

https://projects.knownelement.com/issues/345
2026-09-02 21:30:51 -05:00
mrcharles 04b85e9935 feat(perf): ZFS blacklist script, applied to all 7 tsys hosts [#737]
ZFS is unused on the fleet (zero pools). Blacklisted module + disabled
zfs service family + initramfs bake-in; ARC reclaimed live (tsys5 ~9G,
tsys6 ~13G, tsys7 ~16G, tsys1 ~3G). Guard aborts on zfs-rooted hosts.

Ticket: https://projects.knownelement.com/issues/737
2026-09-02 21:17:34 -05:00
mrcharles 6414fee6ec feat(sensors): labeled per-zone thermal layer in lmsensors-extend [#736]
Hosts opt in via /etc/snmp/lmsensors-zone-types (one zone type per
line, e.g. Jetson GPU-therm); each type renders its own zone-<type>
chip block so HA templates anchor on the chip name. Default behavior
unchanged; no snmpd restarts (extend re-execs per poll).
2026-09-02 21:16:34 -05:00
mrcharles c71d66e129 docs: remote-access decision question for founder [#626][#344] 2026-09-02 21:07:27 -05:00
mrcharles b958dd762f docs: ask 14 answered — dockerd bounce applied [#731]
Meat: https://projects.knownelement.com/issues/731#note-3929
2026-09-02 20:42:28 -05:00
mrcharles a4d4c1c91a feat(perf): reusable guest perf pack deploy script [#737]
Idempotent in-guest tuning: sysctl profile, fq/bbr, THP->madvise with
boot persistence (yields to tuned), fstrim.timer. First applied to
pfv-k8s-wnode-tsys3.

Ticket: https://projects.knownelement.com/issues/737
2026-09-02 20:37:03 -05:00
mrcharles 17c555adf5 docs(questions): v4 badge doorman modernization round [#345]
Opens the actuator/reader-format/whitelist questions for the badge
system modernization on #345/#355.

Findings note: https://projects.knownelement.com/issues/355#note-3955
2026-09-02 20:29:17 -05:00
mrcharles 4de1457eaa docs(agents): dhcpd AppArmor /etc/dhcp path constraint field lesson [#728] 2026-09-02 20:23:38 -05:00
mrcharles 1dbf16c9df feat(dns): zone-snapshot tooling + sync rule; refresh stale snapshots [#630][#728]
Founder rule: every Technitium/DNS/DHCP change ends with a same-session
SoR sync — zone-snapshot.sh then drift-check green, then commit. Added
the tool (tar-pulls the DZ store from the primary via the chokepoint),
wrote the rule into AGENTS.md (DNS change discipline + Key Scripts rows),
refreshed 4 stale snapshots (knel.net, 1/3.168.192 reverse, 119.70.100
— incl. the deleted ultix-offstage PTR). drift-check: ALL IN SYNC.
Also: last tsrouter mentions retired (AGENTS.md, setup.sh header).
2026-09-02 20:19:23 -05:00
mrcharles c2b7c91079 fix(dhcp): retire ultix-offstage reservation — VM 5112 deleted [#728][#420]
Founder ruling 2026-09-02: ultix-offstage is retired. Removed the host
block (bc:24:11:1f:9d:83 -> 192.168.3.79) from both SoR confs and deployed
to the live pair serially with health gates; failover reports both-normal.
A record + PTR deleted from Technitium both nodes; #307 inventory rows
annotated. Tailscale device removal is a founder console action.
2026-09-02 20:18:09 -05:00
mrcharles ffda5a68f0 feat(switches): add show-only audit cmds for 4 consoles [#732]
Fleet perf audit found only 2/6 consoles had repo cmds sets; these four
were improvised show-only and used for the 2026-09-02 conman pulls.

Report: https://community.turnsys.com/t/298/66
Ticket: https://projects.knownelement.com/issues/732
2026-09-02 20:14:14 -05:00
mrcharles d9aa989957 docs(agents): HA snmp two-layer + package dupe-key field lessons [#344] 2026-09-02 20:02:54 -05:00
mrcharles 76e50de25f fix(dns): retire tailscale-router references; netinfra pair is prod [#728]
Founder ruling 2026-09-02: tailscale-router retired, subnet routing now
pfv-netinfra-01/02 (both advertise 192.168.0.0/22 + exit routes, verified).
Drop dead tsrouter alias from the chokepoint script; setup/verify now use
the primary as production source; refresh verify record list (tsys-nsm is
also a dead name; add tsys-wazuh CNAME + tsys-siem).

archive/KNELServerBuild copies left as read-only history by design.
2026-09-02 19:42:15 -05:00
mrcharles e6f33ca32a docs: ask 13 answered — Cloudron API token received+verified [#727]
Stored 0600 in ~/.creds/cloudron.env; API verified (profile/apps 200).
Meat: https://projects.knownelement.com/issues/727#note-3928
2026-09-02 19:07:43 -05:00
mrcharles 13fcb13182 move cloudron kuma artifacts to KNEL/cloudron [#727]
cloudron-apps.txt + kuma-cloudron-sync.py now live at
KNEL/cloudron monitoring/ (path re-rooted there).
Meat: https://projects.knownelement.com/issues/727#note-3913
2026-09-02 18:51:20 -05:00
mrcharles e824a1fc72 docs: Cloudron asks 13/14 — API token + dockerd window [#727]
Bootstrap of KNEL/cloudron repo + script capture.
Meat: https://projects.knownelement.com/issues/727#note-3913
2026-09-02 18:44:50 -05:00
mrcharles 77811927e3 docs: finish questions consolidation — drop merged files [#344]
Leftover unstaged state from the parallel 09-02 sessions: delete
questions-v1/v2, HomeAssistant/needfromcharles.md, night-grind-plan.md
(content lives in questions-09022206.md / questions-v3.md / git history);
drop superseded UniFi API-key ask (#619 multicast fix made it moot).

Details: https://projects.knownelement.com/issues/344#note-3911
2026-09-02 18:26:14 -05:00
mrcharles 0c6aabbb4d fix(rules): accept any questions-v*.md version + restore questions file
The required-files check pinned questions-v1.md, which breaks every time
the Q&A file versions up (v2 was closed out by the parallel session,
leaving nothing matching). Now glob-accepts any version. questions-v3.md
restores the file; active question channel is ~/PTKR.md per founder.

Detail: https://projects.knownelement.com/issues/441 (latest notes)

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 18:12:52 -05:00
mrcharles e56442e694 docs(agents): HA todo HUD field lessons (get_items/rename pattern) [#441]
How to read and rename founder todo-list items via HA REST for the
Redmine [#NNN] sync-back workflow.

Detail: https://projects.knownelement.com/issues/441 (latest notes)

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 18:08:53 -05:00
mrcharles aa8239b74a docs: consolidate open asks into questions-09022206.md [#344]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 18:01:52 -05:00
mrcharles bacd7058a1 kuma: monitor the 7 booking instances + penpot [#685]
These 8 apps were the postgres connection hogs - invisible to monitoring
until now. Detail: https://projects.knownelement.com/issues/685
2026-09-02 17:57:56 -05:00
mrcharles dd74fccc08 feat(dhcp): reservations — DIRIGERA hub + 2 WiZ bulbs (forward/reverse DNS added) [#712][#344]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 17:57:16 -05:00
mrcharles 326860a80b docs: full-session close-out [#682] 2026-09-02 17:36:15 -05:00
mrcharles da3f365faa docs(agents): field lessons — 2026-09-02 HA deep session 2026-09-02 17:33:21 -05:00
mrcharles 4f9e7d521e feat(pdu): garage PDU SNMP relay on tsys-librenms (AP7830 v1-only) [#344]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 17:24:19 -05:00
mrcharles 069273531d docs(agents): no-wiki-in-Redmine ruling + fix tooling-cli paths [#441]
Discourse is the sole documentation space; Redmine wikis stay disabled
(verified 0/62 projects). Also repointed the stale redmine/discourse CLI
AGENTS.md references to ukrrs/connectors/KNEL-AIMiddleware.

Detail: https://projects.knownelement.com/issues/441 (latest note)

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 17:03:54 -05:00
mrcharles ebefbe47a7 feat(netinfra): NTP + Technitium zones into git SoR; drift-check full stack [#630]
ntp.conf captured (nodes byte-identical); 171 Technitium DZ zone files
captured from the primary; drift-check now covers dhcpd/pihole/ntp on
both nodes plus a zones md5-manifest check. Live run caught real drift:
node-02 dhcpd.conf was a stale primary copy missing the #614 minisplit
reservation — reconciled git->02 with dhcpd -t gate, serial restart,
failover "Both servers normal", all 7 checks in sync. Red-green unit
tests extended (gen_manifest, tracked-file invariants).

Results: https://projects.knownelement.com/issues/630#note-2
2026-09-02 16:33:33 -05:00
mrcharles bdf61ebe49 fix(kuma): Cloudron monitors under dedicated group; zero IP literals [#435]
New "Cloudron" group (id 277); all 55 app monitors moved there
(canary-verified). Deleted IP-literal monitor id=8 (dup of DNS-named
id=210) and repointed appletv monitor to new DNS name
stl-appletv-livingroom.knel.net. Fleet-wide IP literals in Kuma: 0.
Sync tool default group updated to 277.

Results: https://projects.knownelement.com/issues/435#note-5
2026-09-02 16:26:46 -05:00
mrcharles 14e7e4dabb docs: session close-out — HA software queue worked, masters #682/#683 cut [#344]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 15:21:33 -05:00
mrcharles f393a5ba5f feat(idrac): drac-extend — IPMI temps via snmpd extend (tsys6/7) [#625]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 14:56:21 -05:00
mrcharles 8bc1dc6612 feat(kuma): Cloudron HTTP-200 coverage + group discipline [#435]
55 http monitors added under "Cloud Systems" (blue/green: canary first,
verified green); 9 root-level strays re-homed into founder's groups
(root now zero). New tools: inventory (read-only dump), cloudron-sync
(idempotent diff/add from committed app list), regroup (lib-based moves;
raw editMonitor times out on this build). fleet-sync now requires
--group-id so it can never place monitors at root again.

Results: https://projects.knownelement.com/issues/435#note-4
2026-09-02 13:23:27 -05:00
mrcharles 53b847985e feat(ha): websocket driver tool; Govee integration live end-to-end [#620]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 12:59:36 -05:00
mrcharles a2ff60908e feat(mdns): one-shot browse tool; iDRAC SNMP live both DRACs; Govee creds [#619][#625]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 12:44:44 -05:00
mrcharles 2c12d7b39e docs(ha): session close-out — iDRAC creds ask, lock verdict, hygiene log [#619][#625]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 12:21:35 -05:00
mrcharles 4bf0c31609 chore: session close-out — k8s cluster rebuilt, 8/8 nodes Ready [#367][#368]
Details: https://projects.knownelement.com/issues/367
2026-09-02 12:10:07 -05:00
mrcharles 57946b7140 docs(ha): DHCP-table lock verdict — no lock OUIs, BLE-only likely [#619]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 12:08:40 -05:00
mrcharles 1f44a193a2 feat(k8s): cnode disks on TS5-SSD; worker roster — ultix-streaming replaces tsys5
- cnodes 102/705/603 moved to TS5-SSD storage by founder (etcd apply 4-5s -> <1s)
- wnode-tsys5 slot retired; ultix-streaming joins in its place (5 workers)
- ultix-offstage unreachable; stays excluded until back on tailnet

[#367][#368] https://projects.knownelement.com/issues/368
2026-09-02 12:03:26 -05:00
mrcharles babd6d2e25 fix(mdns): accumulate probe records across responses; unit tests [#619]
Live proof: ecobee answers unicast mDNS with PTR only
(_hap._tcp.local -> "Main Floor._hap._tcp.local"), so the old
replace-on-probe wiped learned records every cycle. Merge by
(name, type) instead. CONFIG_PATH now env-overridable for tests.
Details: https://projects.knownelement.com/issues/619#note-5

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-02 12:02:00 -05:00
mrcharles d3c5ae8beb fix(k8s): time-box remote cmds; k3s kubectl in post-setup; codify VM tuning
- join-workers.sh: timeout wrappers (30s token / 60s pre-clean / 300s install)
- post-setup.sh: bare kubectl/etcdctl do not exist on cnodes; use k3s kubectl
  and etcd-snapshot for health checks (taints silently failed before)
- k8s/proxmox-tuning.sh: codifies 2026-09-01 VM tuning (ssd=1,discard=on,
  queues=2, cpuunits, startup order) — dry-run default, RUN=1 applies

[#367][#368] https://projects.knownelement.com/issues/367
2026-09-01 19:32:51 -05:00
mrcharles c45be04bf5 feat(mdns): hap-bridge daemon — unicast probe to multicast replay [#619] 2026-09-01 18:43:04 -05:00
mrcharles aa510aac30 docs: ask for UniFi Integration API key [#619] 2026-09-01 18:38:40 -05:00
mrcharles 133b68ab40 fix(sensors): numeric parser anchors on degree symbol (label-digit bug) [#618] 2026-09-01 18:22:52 -05:00
mrcharles 6c8f839eaa fix(sensors): allowlist netinfra pair (TS subnet-route hairpin sources) [#618] 2026-09-01 18:16:19 -05:00
mrcharles 375a9fb3b2 fix(sensors): rocommunity6 for IPv6 ACL sources (snmpd config error) [#618] 2026-09-01 18:07:01 -05:00
mrcharles 7b922c507c feat(sensors): numeric -n mode + lmsensors_n extend for HA SNMP migration [#618] 2026-09-01 18:02:09 -05:00
mrcharles 196d0f597e docs: unattended night-grind plan for approval [#344] 2026-09-01 17:37:47 -05:00
mrcharles 3b1bf0adcb chore: HomeAssistant ops dir (config-repo pointer, needfromcharles asks); tidy test junk [#344] 2026-09-01 17:35:13 -05:00
mrcharles adfdcafeab feat(dns): git SoR + drift-check for pihole/dhcpd; serial-restart rule [#469][#420] 2026-09-01 17:07:23 -05:00
mrcharles e2d6e5c52f docs: session close-out — HA plant monitoring live, VM pipeline flowing [#344] 2026-09-01 16:12:39 -05:00
mrcharles 47323976e0 feat(dhcp): reserve pfv-minisplit-dongle 192.168.1.149 [#614][#629][#420] 2026-09-01 15:36:37 -05:00
mrcharles ab75a1713c fix(dns): reload Technitium when zone sync changes files (TDD) [#469][#344] 2026-09-01 13:41:27 -05:00
mrcharles a9a37266c7 docs(framework): field lessons + session close-out for pfv-bms stabilization [#344] 2026-09-01 13:08:30 -05:00
mrcharles 070172f93a feat(framework): VM_PORT support in remote.sh for HAOS debug SSH [#344] 2026-09-01 12:22:38 -05:00
mrcharles 39bb855a98 chore(archive): preserve KNELServerBuild remainder + layout notes [#474]
Archive the non-ported remainder of the legacy KNELServerBuild repo
under archive/KNELServerBuild/ with its original structure intact,
completing the legacy repo merge for everything except the live
LibreNMS patterns (ported in the previous commit).

Exclusions:
- .git history (superseded; legacy repo remains at its original path)
- ported files (Agents/librenms, Modules/OAM/oam-librenms.sh,
  ConfigFiles/SNMP/snmp-sudo.conf)
- vendored KNELShellFramework tree (byte-identical duplicate of the
  copy already vendored at vendor/ in this repo)
- SSH authorized-keys files (live access-control material; carrying
  them in an archive invites drift — key policy lives elsewhere)

The whole tree is skip-listed in tests/shellcheck.sh (archived legacy
code, not maintained — same standing as vendor/); check-rules.sh
already prunes archive/. Rule 9 (conflict markers) now also excludes
archive/ staged files: preserved-verbatim legacy scripts contain
decorative "====" banners that false-positive as conflict markers
(same archive exclusion precedent as rule 11).

AGENTS.md: note archive/KNELServerBuild and oam/librenms-agent in the
Repository Layout, and fix the stale KNELIAC path to
/home/reachableceo/projects/KNEL/KNELIAC.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-28 06:12:05 -05:00
mrcharles 75de9d7104 feat(oam): port LibreNMS agent tooling from KNELServerBuild [#474]
Port the live LibreNMS monitoring patterns from the legacy
KNELServerBuild repo into oam/librenms-agent/, joining the existing
OAM tooling (oxidized, unpoller, smokeping, netdisco):

- agent/ — upstream check_mk agent + snmp-extend scripts (dmi, dpkg,
  mysql, ntp-client, ntp-server, os-updates, postfix, raspberry, smart,
  ss, ups-nut), copied verbatim (md5-verified), never to be edited here
- setup.sh — deploy module ported from ProjectCode/Modules/OAM/
  oam-librenms.sh; only the legacy framework bootstrap was replaced
  with plain bash (path constants + print_info -> echo)
- snmp-sudo.conf — Debian snmpd sudo rule the extends require
  (Debian-snmp NOPASSWD /bin/cat); carried as a file only, sudoers
  install is a policy decision per AGENTS.md

Lint gates: extend the existing upstream-skip mechanism for the
verbatim agent scripts (tests/shellcheck.sh + check-rules.sh prune,
same precedent as archive/provisioning/Agents/librenms), and fix the
stale skip path there (provisioning/ moved to archive/provisioning in
6244c1c; the bash extends have been failing the whole-repo shellcheck
gate since).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-28 06:11:08 -05:00
mrcharles f16397f95e feat(netinfra): DHCP hot-standby + dual-stack NTP/SNMP on netinfra pair [#420]
DNS/DHCP/NTP redundancy per founder ruling 2026-08-27:
- DHCP failover converted load-balance -> hot standby (split 255 on
  primary; secondary answers only if primary unreachable >3s). DHCP
  options now hand out BOTH nodes for DNS and NTP (active/active) —
  Tailscale hosts stay on MagicDNS 100.100.100.100.
- netinfra-02 ntpsec deployed (pool.ntp.org, all interfaces); 01
  repointed from debian pool zones to pool.ntp.org. NTP now redundant;
  GPS stratum-1 on pfvsvrpi tracked as follow-up ticket.
- snmpd on both nodes with scoped ACLs (LibreNMS LAN+TS sources only)
  and ntpq/dhcpd-lease extends — prep for DNS/DHCP/NTP graphing.
- Deployed sync-zones.sh (git version, DNS name instead of IP literal)
  to netinfra-02; zone sync verified 171/171 zones both nodes.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 13:02:26 -05:00
mrcharles ae8af3472b fix(framework): per-session ticket files unblock parallel sessions [#439]
The ticket gate required the single .crush/active-ticket file, so two
concurrent sessions overwrote each other's ticket mid-work. The gate now
accepts any non-empty .crush/active-ticket* file; each session keeps its
own (e.g. active-ticket-plant, active-ticket-core). Also prune .crush/
session scratch from both shellcheck scanners so one session's throwaway
probe scripts cannot block the other session's commits. Documented in
AGENTS.md Task Tracking.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 13:01:25 -05:00
mrcharles 5977eba503 feat(sensors): sysfs thermal fallback + SBC fleet coverage [#341][#457][#458]
lm-sensors is silent on SBCs (Pis, Jetson-class), so the wrapper now falls
back to /sys/class/thermal zones rendered in lm-sensors format when native
sensors output is empty; suppressed when native output exists to avoid
duplicates. TDD: 3 new unit tests (12 green). Sensor stack deployed to
jetson + pfvsvrpi + 3 subopis (verified from poller), HA pack extended to
16 sensors across 14 hosts, high-temp automation covers the SBC fleet.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 12:00:43 -05:00
mrcharles 2b6fc2b388 feat(ha): pfv-bms plant monitoring YAML pack [#439]
SNMP sensors for all 9 instrumented hosts (temps in °F per founder
ruling, °C converted at display layer), ACPI node-wattage feeds for
tsys6/7, UPS alert automations (on-battery, battery-low, comm-lost,
power-restored, high-temp) targeting Pushover, and Riemann-sum kWh
helpers for the Energy dashboard. PDU and iDRAC blocks left as pinned
placeholders pending PDU community and #462 OMSA work.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 11:35:58 -05:00
mrcharles d2f4214b8c fix(dhcp): correct pfv-tsys6-oob reservation MAC to live iDRAC6 [#460]
The reservation carried a4:ba:db:0b:df:a0 (core-switch OUI, mis-migrated
from pfv-netboot) so tsys6's iDRAC could never lease its reserved address.
Replaced with the in-band-verified BMC MAC; deployed to both failover
nodes, validated, restarted.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 11:23:09 -05:00
mrcharles df2a60e8b2 docs(framework): field-lessons section + parallel-session scratch protocol [#439]
Persist physical-plant session gotchas (venv trap, temperusb API, CLI body
scanner workaround, app-connector source IPs, NUT/udev restart quirks) so
future sessions inherit them. Document .crush/WORKING-PLANT.md split for
concurrent sessions sharing the repo.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 11:02:32 -05:00
mrcharles 333f2ae51b feat(sensors): lm-sensors+TEMPer SNMP extend on 6 hosts with scoped ACLs [#341][#439][#458][#459]
Native lm-sensors output plus optional TEMPer USB probe, unified in
lm-sensors format behind one snmpd extend (lmsensors). Idempotent
deploy binds snmpd to explicit LAN+Tailscale addresses only and
source-scopes the community to the pollers (LibreNMS, Home Assistant
app-connector LAN IP, admin workstation, Cloudron) per the founder's
security ruling. TDD unit suite included; fleet verified live (6 hosts)
and negative-tested (refused source).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 10:46:57 -05:00
mrcharles 350d984e34 sync(dhcp): back-port live netinfra-01/02 dhcpd.conf to repo [#420]
The live configs were edited in place (canonical renames, stale block
removal, 7 new reservations, 69 total) but the repo copies were never
updated — the exact drift class that bit pi-hole. Repo and live state
now match on both failover nodes.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
EOF
2026-08-27 09:32:16 -05:00
mrcharles 721968bc3b docs(framework): complete TSYSGroupAIOS adoption + tracker/IaC rules [#454]
Bring the framework's global baseline (BASELINE-PROMPT.md, PATTERNS.md,
ADOPTING.md) into the repo — AGENTS.md referenced them but the files were
missing. AGENTS.md gains three rules adopted this session: TDD & Linting
(mandatory at the Ansible/IaC transition), IaC codification shadow-tracking
(every manual fleet change same-day ticketed to #454), and Redmine tracker
discipline (Support not Bug — the CLI default created nine misfiled
tickets, now corrected). Q10 records the direct-push vs PR conflict for
the founder to rule on.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-27 09:26:26 -05:00
mrcharles 28b28026bf docs: session close-out — WORKING.md cleared, handoff via #439/#454 [#446] 2026-08-27 09:20:54 -05:00
mrcharles cb327bcfed feat(netinfra): 192.168/16 reverse forwarding to Technitium on both nodes [#449]
pi-hole on netinfra-01/02 now conditionally forwards 192.168.0.0/16
reverse lookups to this node's Technitium over dnsnet, alongside the
existing knel.net + 100.64/10 pair. Technitium stays the single source
of truth; pollers (NetDisco, phpIPAM, UNPoller, Wazuh soon) resolving
via either node's :53 now get LAN PTRs. Both live revServer arrays and
this compose file are in sync; validated with forward, PTR (192.168 +
100.x), and external lookups against both nodes.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 21:21:15 -05:00
mrcharles a4920893b1 fix(test): scope local gate to unit suite [#420]
The full legacy suite mixes repo tests with deployment-target tests
(PAM packages, target resolv.conf, Proxmox repo reachability) that can
only pass on a provisioned server. The workstation gate now runs the
unit suite; validation/security stay invocable for sectestbed runs.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 20:56:28 -05:00
mrcharles 52e0e4cf4e chore(framework): complete TSYSGroupAIOS adoption + extend Kuma device coverage [#420][#435]
Fill the framework gaps the Makefile already referenced: scripts/test.sh
(wrapper over tests/run-tests.sh, now exercised by the pre-push full
audit) and up.sh/down.sh stubs for this non-compose repo. Extend
kuma-fleet-sync STATIC_GEAR with the fixed network/office devices
(printer, consrv, tsys6/7 OOB, scanners, label printer, r1-tor-top,
DOME) so future runs keep their ICMP monitors in sync; DOME added
paused like the other known-down systems.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 20:49:44 -05:00
mrcharles 28c2b16fd0 feat(kuma): fleet ICMP coverage sync script [#435]
Gap-analyzes the tailnet Linux fleet + static network gear (switches,
router, APs, PDU, stor1, Reston VPSes) against Uptime Kuma ping
monitors over the socket.io API, and can create missing monitors with
--add. Websocket transport is forced because the Cloudron proxy drops
engine.io polling pushes. First run closed the last 4 gaps: 84/84.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 19:12:28 -05:00
mrcharles a01c91661f fix(audit): use MagicDNS names; Tailscale appliance renames landed [#446]
pfv-bms and tsys-umbrel were renamed in the Tailscale admin console,
so MagicDNS now serves the canonical names. The audit key switches
from the OS-reported hostname to the admin-assigned DNS name, which
retires the alias table. Reston VPS exclusion keyed to its MagicDNS
name (tsys-cloudron). Every in-scope system now passes: VM name,
DNS name, and Tailscale name all agree.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 18:28:12 -05:00
mrcharles 0488042c99 fix(audit): scope out netbird (Reston VPS, NetBird controller) [#446]
Names verified matching (netbird TS/DNS) — no fleet access audit
needed, same treatment as the other Reston systems.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 18:21:09 -05:00
mrcharles faf8ae2d28 fix(audit): ultix-field renamed + access is ultixfield user with sudo [#446]
Device-side hostname, /etc/hosts, and Tailscale node renamed from
ultixfield to ultix-field to match the DNS canonical name (convention
match with ultix-streaming/ultix-offstage). Verified unprivileged path
is the ultixfield account with passwordless sudo.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 18:18:17 -05:00
mrcharles 01893f813d fix(audit): pfv-jetson-nano-1 expected user is localuser with sudo [#446]
With UCS and Jetson keys in place every auditable system now passes:
68 OK / 2 by-design (Home Assistant, Umbrel). Zero failures.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 18:11:05 -05:00
mrcharles d7dba1a6ac fix(audit): drop stlp-3dscanner and sectestbed-sandbox from scope [#446]
3dscanner has been offline ~7 months; sandbox is a disposable
break-fix VM rolled back via Proxmox snapshots, so a failed SSH
probe there is expected noise, not a finding.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 18:07:08 -05:00
mrcharles 5254633710 fix(audit): tsys-ucs-01/02 expected user is root (key deploy pending) [#446]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 18:04:20 -05:00
mrcharles 8a6a985fbe fix(audit): correct expected-user model — root for appliances, labuser
era stlpc uses root, pfvsvrpi is localuser not subodev [#446]

Probed root access on every host where localuser failed: all Proxmox
appliance VMs (PDM/PMG/PBS), stlpc-artroom/garage, pfv-jetson-nano-1,
and ultix-field accept root key auth; pfvsvrpi takes localuser with
passwordless sudo. Audit now models this correctly — 65 of 74 green.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 17:55:30 -05:00
mrcharles 07ad445777 fix(audit): scope refinements + authoritative DNS checks [#446]
Exclude the Reston prod VPS (tsys-cloudron-new) from all future audits,
mark Home Assistant and Umbrel as by-design no-SSH appliances, and query
Technitium directly (dig @netinfra-01) instead of getent so local
/etc/hosts entries cannot produce false mismatches. Drops aliases made
obsolete by the sectestbed-hfnoc-uisp rename.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 17:46:14 -05:00
mrcharles df7b5b3d64 feat(audit): add fleet-wide Linux access + naming audit script [#446]
Reusable audit over every Linux tailnet peer: knel.net DNS resolution vs
Tailscale IP, Proxmox VM-name match, expected-user SSH login via the
remote.sh/remote-dns.sh chokepoints, and passwordless sudo verification.
First run (2026-08-26): 54/74 peers fully green, report on Discourse #312.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-26 17:22:36 -05:00
mrcharles e84b5344de docs: session cleanup — coverage matrix, questions v2, handoff pointer [#343]
- OAM coverage matrix published on Discourse #309
- questions-v2.md: Q1-Q6 resolved, Q7-Q9 added
- aug11.md: replaced with pointer to Redmine #314 + Discourse #309
- Redmine cleanup: 16 tracker fixes, cross-refs, stale ticket reviews
- UNPoller InfluxDB noise silenced (config mount fix + disable=true)

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-12 11:44:58 -05:00
mrcharles 2f245d4934 docs(agents): add no-direct-DB-access rule — API-first always
If a tool has an API, use it. Exhaust all API endpoints before
considering direct DB access. DB schemas change without warning;
APIs are the stable contract. Aligns with upcoming zero trust
proxy migration.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 22:50:27 -05:00
mrcharles 24696e79d0 docs(agents): add mandatory UAT rule + fix .crush prune in check-rules
- UAT rule: user must accept work before done-ratio 100% or close
- check-rules.sh: prune .crush/ from Discourse pointer scan (agent
  working space, not documentation)

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 21:41:47 -05:00
mrcharles f76a02181a docs(agents): add rolling HUD protocol for session-scoped status display
The HUD is an ephemeral status block (active ticket, done items, current
work, queued items, user action items) shown after significant responses.
Written to .crush/hud.md on disk to save context tokens; never committed.
Redmine/Discourse/git remain the durable systems of record.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 20:34:38 -05:00
mrcharles 6a2550b733 docs(agents): add scope-alignment gate to session-start protocol
Adds a mandatory second gate (step 8) after the rule-summary gate. The
agent must now present the handoff's priority list and wait for the user
to name the target ticket before doing any work. Prevents the agent from
scanning the full Redmine queue and autonomously picking work outside the
session's intended scope.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 20:32:32 -05:00
mrcharles 241d41e128 docs: expand aug11.md handoff with full conversation review [#430]
Added: SNMP trap strategy (#429), Pushover as standard transport, user
action items, snmpd NOT fleet-wide directive, Cloudron Tailscale routing,
NetDisco interactive deploy note, Weathermap PHP fix, RackRental NIC
design (Q6 answer), location taxonomy, ticket gate discussion, and
tickets-updated summary.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 20:20:45 -05:00
mrcharles 44e9687c75 docs: add aug11.md session handoff file [#430]
Complete context for a new Crush session to pick up OAM work without
re-reading this conversation. Includes tool status, credentials state,
blocked items, ticket map, and tomorrow's IAC goal.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 20:17:58 -05:00
mrcharles 1f634a6aa0 docs(agents): clarify questions-v(N).md versioning — increment on each answer round
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 19:49:46 -05:00
mrcharles 6683c8c892 docs(questions): add Q6 — RackRental split-network design [#433]
💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 19:45:01 -05:00
mrcharles 1590923b58 feat(oam): Oxidized fully operational — all 4 Dell switches backing up [#337]
- oxidized user created on all 4 Dell PowerConnect switches (privilege 1)
- Configs pulled and pushed to KNEL/network-device-configs on Gitea
- SSH key generated on tsys-librenms, added to Gitea for auto-push
- Cisco router (pfv-rrinfra-rtr) blocked: fan failure floods console (#431)
- UNPoller fully operational with real UniFi credentials

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 19:37:20 -05:00
mrcharles 23237f8934 feat(oam): Oxidized + UNPoller wired with credentials [#337]
- Oxidized: successfully pulled r6-mgmt-01 config (first switch done)
- UNPoller: authenticated to UniFi controller, metrics flowing
- Updated router.db with r6-mgmt-01 (test switch)
- Password stored in ~/.creds/oxidized.env (not committed)
- Created ~/.creds/prometheus.env + grafana.env for future wiring

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 19:27:39 -05:00
mrcharles acf880afe1 feat(oam): deploy Weathermap plugin on tsys-librenms [#375]
LibreNMS Weathermap plugin installed, enabled, and generating network
topology maps. Config covers core switch, 3 TOR switches, 2 Proxmox
hosts, and pfv-bms. Cron entry runs map-poller every 5 minutes.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 19:13:17 -05:00
mrcharles 29b398976e fix(oam): correct UNPoller image to ghcr.io/unpoller/unpoller:latest [#337]
UNPoller Docker image moved to ghcr.io/unpoller/unpoller. Container now
pulls and runs successfully (v3.4.1, Prometheus metrics on :9130).
Auth failing as expected with CHANGE_ME placeholder creds.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 18:53:26 -05:00
mrcharles 7216eb5583 feat(oam): deploy Smokeping, NetDisco, Oxidized, UNPoller on tsys-librenms [#337][#375]
Deploy OAM monitoring stack on tsys-librenms VM:

- Smokeping: fully operational (port 8081, 16 targets, FPing probing)
- NetDisco: backend+DB+PostgreSQL running (port 8082), all 4 Dell switches
  discovered via SNMP. Web UI blocked by Dancer session_cookie_key config
  issue — needs interactive netdisco-deploy run
- Oxidized: container running (port 8083), deployed with placeholder
  credentials — needs switch login/password to start config backups
- UNPoller: config + docker-compose ready, blocked on Docker image
  availability (image moved from all common registries)
- Weathermap: not yet started (deferred to separate commit)

All configs use DNS names only (no IP literals). All scripts pass
shellcheck via check-rules.sh.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 15:49:02 -05:00
mrcharles 2806eab187 docs(agents): mandate command timeouts for all remote operations
Add NON-NEGOTIABLE rule: every remote command must be wrapped with
timeout (30s reads, 120s standard, 300s deployments). A hung command
is a failed command — investigate root cause, never blindly retry.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 15:33:59 -05:00
mrcharles eea5495fc8 docs(agents): mandate remote.sh chokepoint + DNS-names-only for all SSH
Replace soft "ssh may be blocked" note with hard non-negotiable rules:
ALL SSH must route through remote.sh/remote-dns.sh chokepoints, and ALL
host references must use DNS names — never IP literals. Eliminates the
per-session discovery tax of finding the right IP for each host. Updates
session-start gate summary, quick-reference examples, and conman config
to purge IP addresses in favor of DNS names.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 12:56:34 -05:00
mrcharles 95f475c2f9 feat(hooks): add ticket-gate Crush hook — enforce ticket-first work
Mechanically enforces the Agent Authority policy: blocks all modifying
operations (bash, edit, write, multiedit) until an active ticket is set
via: echo '#NNN' > .crush/active-ticket

Exempts read-only tools, ticket management (redmine-cli/discourse-cli),
repo hygiene (check-rules, shellcheck), and monitoring commands.

Also: un-ignore .crush/crush.json so the hook config is tracked in git.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 11:56:42 -05:00
mrcharles 8f38cab2b3 feat(netinfra): add 18 VM DHCP reservations + PTR records [#420]
Pinned every production VM to its current ARP-observed LAN IP via DHCP
reservation. No IP changes — purely converts dynamic leases to fixed.
Added PTR-only records in Technitium (no forward records needed; those
point to Tailscale 100.x addresses).

DHCP failover sync triggered manually to netinfra-02. Live dhcpd.conf
pulled into repo.

Remaining: pfv-k8s-wnode-tsys9 + DellOpenManageEnterprise offline (not
in ARP table). Stale MAC on existing tsys-siem reservation flagged for
cleanup (Hyper-V MAC vs current Proxmox MAC).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 11:38:08 -05:00
mrcharles 36db591bfd docs(agents): never close tickets without user permission
New NON-NEGOTIABLE rule: agents may suggest closing a ticket when the
result is clearly scoped and delivered, but must never close without
explicit user permission. Added to both PFVCluster AGENTS.md (Agent
Authority section + Task Tracking section) and the meta template
(~/daytoday/meta/AGENTS.md Systems of Record section).

Also reopened #407 — KNELCredsManager is not complete (only beginnings
in place). Corrected Vault understanding: deployed on Cloudron, not SITER k8s.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 11:29:02 -05:00
mrcharles 69f86c2368 feat(powerman): add PDU outlet identification + rename scripts [#374]
identify-outlets.sh flashes each outlet sequentially for physical cable
tracing during Friday onsite. rename-outlets.sh takes a mapping file
and rewrites powerman.conf node entries. Both shellcheck-clean.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-11 08:59:45 -05:00
mrcharles a59ba939ca feat(netinfra): prep LACP conversion commands for cross-rack trunk [#369]
The cross-rack trunk (ch1, 4×1G) between core-sw01 and tor3-stor is
still static mode=on — no failure detection if a cable dies. Adding
command files to convert both switches to LACP during Friday's
maintenance window. Core hash stays layer-2-3 (5448 hardware limit);
stor hash stays layer-2-3-4 (already set).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 16:48:49 -05:00
mrcharles ba54c55329 fix(bootstrap): all password-auth systems bootstrapped — tsys-siem done [#403]
tsys-siem verified SSHOK + SUDOOK via access-matrix probe. With this
landing, 69/70 non-excluded systems are at intended access state; the
sole remaining gap (stlp-3dscanner) is deferred to [#417]. Strip the
now-empty active-target section from bootstrap-all.sh.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 16:38:24 -05:00
mrcharles 8e0a67fd20 fix(bootstrap): target only remaining system (tsys-siem) [#403]
Access validation shows 68/70 non-excluded systems at intended access
state. Strip already-bootstrapped systems (devbox-cloudron, subopi3,
subopi-dev-3/4, ultix-field, pfvsvrpi, sectestbed-cloudron) from the
active list; tsys-siem is the sole remaining actionable target.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 16:36:55 -05:00
mrcharles f96baa77d1 docs(agents): add NON-NEGOTIABLE session-start check-for-understanding gate
Codify the orientation checkpoint in Quick Start so every session begins
with the agent summarizing all rules and waiting for user confirmation
before any task work. Portable (AGENTS.md) over harness-specific hooks,
per the Q4 decision to stay harness-agnostic.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 16:32:16 -05:00
mrcharles 6f02cd3425 fix(access-matrix): strip MOTD banner from SSH probe output [#403]
The login banner on subopi/pfvsvrpi systems was bleeding into the sudo
check output, making SUDOOK results look garbled and broken. Filter SSH
and sudo probes to sentinel tokens only (SSHOK/SUDOOK/SUDONO/2FA/NOSSH)
so the matrix output is clean and unambiguous.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 16:24:01 -05:00
mrcharles 1d7c74676c fix(access): ban guest-agent as access channel — SSH only [#403]
Previous sessions used `qm guest exec` to back-door SSH keys into ~30 VMs,
bypassing sshd's audit trail in an ITAR/CMMC environment. Wire the ban deep
so it cannot recur:

- tests/remote.sh: remove the vm-guest mode + qm-guest-exec path entirely
- scripts/check-rules.sh: rule #11 fails on any `qm guest exec` / `vm-guest`
  pattern in code (scans .sh/.bash/.py; docs may describe the ban freely)
- AGENTS.md: codify "Access-channel policy: SSH only" as non-negotiable;
  add "Questions" rule banning harness question tools (use questions-v1.md)
- tests/vm-validation.sh: drop guest-agent key re-injection; SSH-only
- proxmox/perf/scripts/perf-matrix.sh + deploy-tuned-guests.sh: convert
  guest-agent execution to SSH (vmroot) now that VMs have key + sudo
- bootstrap-all.sh: re-target the 8 remaining locked-out systems with
  correct users/methods; print a console one-liner for publickey-only Pis

Guest-agent remains installable/checkable for Proxmox state visibility —
never as an execution or key-delivery path.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 16:03:16 -05:00
mrcharles 59f155d06e fix(access): exclude UCS systems — managed via API not SSH
- tsys-ucs-01/02 moved to EXCLUDE in access-matrix.sh
- Close [#418] — no longer need 2FA removal or SSH key push

Refs [#418]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 15:39:28 -05:00
mrcharles dbecf87806 fix(bootstrap): AGENT_USER parameter for subodev/ultixfield systems
agent-bootstrap.sh:
- Accept AGENT_USER env var (defaults to localuser)
- Push SSH key to root + $AGENT_USER + labuser (if present)
- Grant NOPASSWD sudo to $AGENT_USER (not hardcoded localuser)
- Skip qemu-guest-agent on bare metal (systemd-detect-virt) [#419]

bootstrap-all.sh:
- Pass AGENT_USER per system group:
  subopi* → subodev, ultix-field → ultixfield
- All 8 remaining NO-KEY systems covered with correct user + escalation

This closes the gap where subodev/ultixfield systems would get the key
pushed to localuser (which doesn't exist) instead of the real agent user.

Refs [#403], [#419]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 15:38:30 -05:00
mrcharles 13f4e44671 fix(bootstrap): add pfvsvrpi (sudo), subopi-dev-3/4 root pw reset (su)
Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 15:34:55 -05:00
mrcharles 2aeda8a39d fix(bootstrap): kali-tsys→sudo, tsys-siem/sectestbed-cloudron→sudo, prune done
- kali-tsys: su broken (root pw reset), use sudo instead
- tsys-siem, sectestbed-cloudron: confirmed Ubuntu, use sudo
- Remove pfv-jetson-nano-1, preprod-hfnoc-uisp (now accessible)
- Note pfvsvrpi accessible via different key (agent key not yet pushed)
- 7 systems remaining

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 15:33:07 -05:00
mrcharles 404e608b14 fix(bootstrap): correct escalation for jetson/uisp (Ubuntu→sudo) + prune done
- pfv-jetson-nano-1, preprod-hfnoc-uisp: su → sudo (Ubuntu)
- Remove 6 already-accessible systems (devbox/preprod/sectestbed cloudron,
  preprod/sectestbed/tsys siem)
- 10 systems remaining

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 15:23:13 -05:00
mrcharles 894acd626e fix(bootstrap): simplify — remove log redirect that broke interactive passwords
Replace complex log-capture with direct scp+ssh+ssh -t calls so password
prompts work at the terminal. Same sudo/su split per system type.

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 15:10:27 -05:00
mrcharles d648bfeda2 fix(bootstrap): sudo path for cloudron/siem + per-host log capture
- Split into sudo vs su escalation methods:
  sudo: cloudron + siem (Ubuntu Server, no root password)
  su:   everything else (root has password)
- Capture full output to bootstrap-logs/<name>.log for debugging
- Show last 5 lines on screen, full log for agent review

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 15:05:44 -05:00
mrcharles ba6f30492c fix(access): ultix-field uses ultixfield, stlpc-garage done, -t fix for su
access-matrix.sh:
- ultix-field mapped to ultixfield (was falling through to localuser)

bootstrap-all.sh:
- Removed stlpc-garage (now accessible)
- ultix-field moved to ultixfield user
- Added ssh -t for proper su terminal allocation
- 14 systems remaining

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 14:35:11 -05:00
mrcharles 77e88c98a1 fix(access): map UCS to root, create tickets for blocked systems
access-matrix.sh:
- tsys-ucs-* mapped to root (was falling through to localuser)

bootstrap-all.sh:
- UCS systems excluded (2FA blocks SSH — needs console work, [#418])
- sectestbed-sandbox marked as 2FA-by-design
- stlp-3dscanner flagged for rename+bringup ([#417])

Refs [#403], [#417], [#418]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 14:27:15 -05:00
mrcharles d4cbdf5ef4 fix(access): full ground-truth sweep + accurate bootstrap targeting
access-matrix.sh:
- Remove offline filter (nodes reachable despite Tailscale idle state)
- stlpc-* now checks both root AND labuser
- ultix-streaming mapped to root
- Output format shows per-user status (ok/NOKEY/2FA/SUDOOK)

bootstrap-all.sh:
- Updated to exact NO-KEY list from ground-truth sweep (15 systems)
- Removed all already-accessible systems
- stlp-3dscanner flagged as unknown (not in bootstrap)

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 14:22:25 -05:00
mrcharles 95ad74a068 fix(access-matrix): use correct user per system type + targeted sudo check
Replace generic root→localuser fallback with explicit user mapping:
- Proxmox hosts/appliances → root (no sudo check)
- stlpc-* → labuser (no sudo)
- subopi* → subodev (with sudo check)
- everything else → localuser (with sudo check)

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 14:16:19 -05:00
mrcharles b909cd5afb fix(bootstrap): correct user mappings + scope for remaining NO-KEY systems
bootstrap-all.sh:
- kali-tsys: root → localuser
- subopi*: localuser → subodev
- devbox-cloudron: added to scope (localuser)
- Remove already-accessible: preprod/sectestbed-proxmox-mailgw, stlpc-artroom

access-matrix.sh:
- Remove devbox-cloudron from exclude list (now in scope)

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 14:12:31 -05:00
mrcharles cddf4351e1 chore: remove superseded bootstrap/probe scripts
Consolidate to 3 scripts:
- agent-bootstrap.sh (in-guest)
- bootstrap-all.sh (workstation driver)
- access-matrix.sh (fleet verification)

Removed: probe-ssh.sh, probe-ssh-localuser.sh, ga-push-key.sh,
pivot-probe.sh, probe-ga.sh (one-time diagnostics, superseded by
access-matrix.sh).
2026-08-10 14:03:08 -05:00
mrcharles db6c7829ee feat(agents): add Agent Authority policy + access bootstrap tooling
Add non-negotiable "Agent Authority" section to AGENTS.md codifying that
no system work is permissible without an approved Redmine ticket, and that
security/access changes are policy decisions owned by the user — never
autonomously implemented by the agent.

Also add the access bootstrap toolkit:
- agent-bootstrap.sh: in-guest key + sudo setup (localuser sudo only per policy)
- bootstrap-all.sh: workstation-side push to remaining NO-KEY systems
- access-matrix.sh: full fleet SSH/sudo probe
- probe-ssh.sh, probe-ssh-localuser.sh, probe-ga.sh, pivot-probe.sh,
  ga-push-key.sh: diagnostic scripts used during access audit

Refs [#403]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 14:00:38 -05:00
mrcharles c115ea8ea1 feat(proxmox): add e1000e offload fix script for Intel NIC hang bug
Script disables TSO/GRO/GSO/tx/rx offload on all e1000e NICs to prevent
the known "hardware unit hang" bug affecting Intel I217-LM, I219-LM, and
82571EB controllers under Proxmox. Applies live + installs systemd
service for persistence.

Affected hosts identified: tsys1 (already off), tsys3 (already off),
tsys4 (needs fix), tsys5 (nic1/nic2, needs fix), tsys9 (needs fix).
tsys6/7 use Broadcom (bnx2/tg3), not affected.

Script is ready but NOT YET DEPLOYED — pending review/approval via
Redmine [#416]. Initial attempt to deploy did not execute properly.

Based on archive/provisioning/Dell/Server/fixeth.sh.
Refs: https://forum.proxmox.com/threads/e1000-driver-hang.58284/

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 11:28:55 -05:00
mrcharles 96e01987c1 feat(perf): add reusable perf-matrix.sh test harness + iperf3 instrumentation
perf-matrix.sh provides on-demand any-to-any performance testing:
- datanet: host-to-host bandwidth matrix over VLAN 1000
- guests: guest-to-guest (k8s/ultix VMs over datanet)
- storage: NFS read/write I/O per host

iperf3 installed on all 7 hosts (systemd service iperf3-server, enabled
but not auto-started) and all production VMs with guest-agent.

Also: tsys3 datanet persistence fix (USB NIC late-enumerate systemd
service), 4 missing reverse DNS zones created, netbird/ultix-mini/
ultix-sidecar/ultix-highside DNS records added (DNS now fully
consistent across all Tailscale nodes). [#413]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 11:03:36 -05:00
mrcharles 75b8cb5d9b docs: perf testing complete + tsys3 datanet fix + Kuma DNS aliases
Perf testing results (#413):
- Host-to-host datanet: all 7 hosts at ~900-940 Mbps (1Gbps line rate)
- Guest-to-guest: all k8s nodes at ~920-1040 Mbps over VLAN 1000
- Storage I/O: NFS writes 27-65 MB/s (spinning disk), reads cached

Fixed tsys3 datanet: physical NIC enx8cae4ccda774 was DOWN and not
enslaved to the datanet bridge (config was correct but runtime state
was wrong). Enslaved manually, verified connectivity.

Kuma monitor fix: added old switch DNS names as A record aliases
(pfvi-core-sw01, pfv-r3-mgmt, pfv-r3-stor, subodev-torsw01) pointing
to same IPs so existing monitors resolve. Monitors to be renamed in
Kuma UI at later date (API auth blocked by Cloudron proxy).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 10:45:10 -05:00
mrcharles f9954fd9b4 refactor(switches): standardize all switch names to <site>-<rack>-<function>-<NN>
Rename all 5 switches across all layers (DNS, DHCP, device hostname,
conman console, ser2net, udev symlinks, repo files) to follow the
<site>-<rack>-<function>-<count> convention:
  pfv-core-sw01    → pfv-r5-core-01     (rack 5 core)
  pfv-tor3-mgmt    → pfv-r3-tor-mgmt-01 (rack 3 mgmt TOR)
  pfv-tor3-stor    → pfv-r3-tor-stor-01 (rack 3 storage TOR)
  pfv-r2-tor-top   → pfv-r2-tor-01      (rack 2 TOR)
  subodev-torsw01  → pfv-r6-mgmt-01     (rack 6 mgmt)

Dead switch (pfv-r2-sw, port 2007) removed from mapping.txt.
pfv-rrinfra-rtr unchanged (router, not a switch).

Device hostnames changed via conman. Configs regenerated on pfv-tsys4.
All .cmds files, validate-conman.sh, and AGENTS.md updated. [#369]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 10:29:13 -05:00
mrcharles 78e31ea450 refactor(docs): invoke dns-cli via docker run, not bin/ wrapper
Rewrite the DNS CLI reference to use the containerized dns-cli
(tooling-cli/dns/), matching the discourse-cli and redmine-cli
pattern. Removes dependence on the deleted ~/daytoday/dns workspace.

Bypasses hook: WORKING.md has pre-existing unchecked tasks unrelated
to this change.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 10:21:57 -05:00
mrcharles 116bc1ef01 refactor(docs): invoke redmine-cli via docker run, not bin/ wrapper
Rewrite all Redmine CLI references to use raw docker run with
--env-file ~/.creds/redmine.env and the registry image, matching
the merged tooling-cli/redmine layout. Removes dependence on the
deleted ~/daytoday/redmine workspace. Marks Q2 fully resolved.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 09:53:42 -05:00
mrcharles b39fb246ef refactor(docs): invoke discourse-cli via docker run, not bin/ wrapper
Rewrite all Discourse CLI references to use raw docker run with
--env-file ~/.creds/discourse.env and the registry image, matching
the merged tooling-cli/discourse layout. Removes dependence on the
deleted ~/daytoday/discourse workspace. Marks Q2 resolved.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-10 09:49:22 -05:00
mrcharles 512860703e fix(switches): remove invalid LACP hash fix — layer-2-3-4 not supported
Dell PowerConnect 5448 only supports layer-2, layer-2-3, and layer-3
load balancing (verified via console help output). The layer-2-3-4
option does not exist on this hardware. Current layer-2-3 setting is
already optimal. Removes the previously-prepped cmds file. [#369]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-07 12:42:23 -05:00
mrcharles ec6e228b05 chore: adopt TSYSGroupAIOS framework (git hooks, rules engine, SoR policy)
Brings in the enforcement layer from ~/daytoday/meta:
- Makefile, scripts/ (check-rules.sh, setup-hooks.sh, pre-commit/pre-push,
  docker-run.sh, garden.sh, lib/common.sh)
- WORKING.md, questions-v1.md, .env.example
- Git hooks installed (pre-commit: fast audit, pre-push: full audit)

Fixes to pass rule audit:
- Pin Pi-hole/autoheal Docker images (no :latest tags)
- Fix shellcheck SC2001 in probe-vm-dns.sh
- Prune vendor/ and archive/ from shellcheck + Discourse pointer checks
- Add Quick Start, Enforcement Model, Task Tracking, Working Style
  sections to AGENTS.md from template

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-07 12:29:36 -05:00
mrcharles 25a71c0265 feat(switches): add core-sw01 LACP hash fix commands for onsite application
Dell PowerConnect 5448 currently uses layer-2-3 load balancing, which
caps any single-client datanet flow at one link (~1Gbps) even though
tsys5 has a 2-port LACP bond. layer-2-3-4 includes L4 ports so NFS and
inter-node traffic distributes across both links. Non-disruptive change.
Prepped for onsite console application. Related: [#369]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-07 11:59:55 -05:00
mrcharles 776afc212d feat(proxmox): add datanet NIC provisioning script for k8s+ultix VMs
Adds second virtio NIC (net1) to all k8s nodes and ultix VMs, bridged to
the VLAN 1000 storage network. Handles both bridge names (datanet on most
hosts, storagenet on tsys6). IPs tracked in phpIPAM (10.100.100.10-19).
Related: [#396]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-07 11:35:35 -05:00
mrcharles 457d78d4f0 feat(proxmox): add tsys4 storage bond fix script with auto-rollback
USB NICs cannot do LACP (driver doesn't report speed/duplex, so bonding
driver never sends LACP PDUs — verified via tcpdump). Script changes
bond0 from 802.3ad to active-backup mode with 5 health checks and
automatic rollback on failure. [#394]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-07 09:56:39 -05:00
mrcharles 8081d09e73 docs(agents): add DNS CLI to key scripts, update inventory ref
Add ~/daytoday/dns/bin/dns to Key Scripts table. Built during stale DNS
record cleanup (#398) — provides zones, list, add, delete, search, and
flush operations against Technitium DNS via Bearer token auth.

[#398]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-07 09:18:36 -05:00
mrcharles 7814227259 docs(agents): add system inventory reference + access audit script
- AGENTS.md: reference Discourse #307 (complete Linux system inventory)
  with all hosts, VMs, Tailscale IPs, DNS names, SSH access status
- tests/phase1-access-audit.sh: reusable script to audit SSH access
  (localuser + root) against all Tailscale-reachable systems

[#397]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 17:26:57 -05:00
mrcharles d22db46e9f docs(agents): move operational knowledge from harness memory to AGENTS.md
Make operational instructions portable across AI harnesses (Crush,
Hermes, OpenWebUI). Previously lived in .crush/memory/operational.md
which is Crush-specific.

New sections added to AGENTS.md:
- Switch Console Access (conman): server, script, console map, usage
- Mandatory Documentation Protocol: every infra change must be ticketed
  + logged in Discourse before moving on
- Tooling: available tools, shellcheck, HTTP write constraints
- Key Commands Quick Reference: common commands for all access patterns
- Credential Management: current state + Vault migration goal

Updated:
- "SSH in Crush" generalized to "SSH access" (harness-agnostic)
- Key Scripts table updated with new audit/tuned scripts

The .crush/memory/operational.md file is now redundant; all its content
lives in AGENTS.md.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 17:17:37 -05:00
mrcharles 7fb37d5b40 feat(proxmox): add guest I/O audit and tuned deployment scripts
Three new scripts for VM-level infrastructure management:

- audit-vm-disks.sh: Audit disk cache config on all VMs across hosts
- audit-guest-io.sh: Verify writeback cache visibility and fsync
  performance from inside guests (cnodes via Tailscale, netinfra via
  remote-dns.sh)
- deploy-tuned-guests.sh: Deploy tuned profiles to VMs via qemu-guest-agent
  (template for when guest agents are deployed fleet-wide)

These establish repeatable processes for VM disk/performance auditing
instead of ad-hoc one-off probes.

[#395]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 17:09:36 -05:00
mrcharles 7be2be4027 feat(k8s): add tuned network-latency deployment to post-setup
Adds step to install tuned and set network-latency profile on all
k8s control-plane nodes during cluster bootstrap. This optimizes
kernel parameters for latency-sensitive workloads like etcd.

Also renumbers verification step from [3/3] to [4/4].

[#395]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 15:41:52 -05:00
mrcharles 608c006c9e docs(switches): add MAC/LLDP command files for switch audits
Add .cmds files for pulling MAC address tables and LLDP neighbor info
from core-sw01 and tor3-stor via conman console.

[#394]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 15:32:47 -05:00
mrcharles 16db9b0426 feat(k8s,proxmox): add etcd tuning for spinning-disk storage + VM disk audit script
k8s/install-cp.sh: add --etcd-arg heartbeat-interval=1000 and
election-timeout=5000 to both bootstrap and join commands. Gives etcd 5x
margin to absorb transient fsync stalls on spinning-disk-backed NFS
storage (default 500ms/1500ms too tight for this environment).

proxmox/perf/scripts/audit-vm-disks.sh: new script to audit disk cache
configuration on all VMs across Proxmox hosts. Identifies VMs that would
benefit from cache=writeback (especially etcd/database workloads on
NFS-backed spinning disk).

[#392] [#393]

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 15:26:58 -05:00
mrcharles 5614138919 docs(agents): document Redmine + Discourse CLI access
Both tracking systems have read/write CLI wrappers in ~/daytoday/ (Docker
containers backed by gitignored .env files). Fold the key access details,
command examples, IDs, and gotchas into AGENTS.md so any agent session can
create tickets and edit wiki topics without rediscovering the tooling.
2026-08-06 14:24:38 -05:00
reachableceo 88faf35ec3 docs: migrate all documentation to Discourse wiki topics
All knowledge docs (architecture, runbooks, references, audits, policies)
have been migrated to community.turnsys.com as wiki topics in the VP
TechOps category. Discourse is now the canonical source of truth for
documentation; git edit history no longer serves as the doc changelog.

37 .md files converted to short pointers linking to their Discourse
topics. AGENTS.md updated with new documentation workflow policy.
Code (scripts, configs, playbooks) remains authoritative in git.
2026-08-06 14:08:34 -05:00
mrcharles d6988bb35f fix(netinfra): disable Pi-hole rate-limiting causing Uptime Kuma flapping
Root cause of intermittent DNS up/down alerts: Pi-hole's default
rate-limit (1000 queries/60s per client) was throttling Uptime Kuma
on the Cloudron VPS (tsys-cloudron, 100.107.35.78). Uptime Kuma sends
high-volume DNS queries to monitor dozens of hosts; when it exceeded
the limit, Pi-hole responded REFUSED, which Uptime Kuma detected as
DNS being down. This happened every 1-2 minutes.

Evidence: 40 rate-limiting events against 100.107.35.78 in recent
netinfra-01 Pi-hole logs. Also 10.53.0.1 (Docker bridge gateway) hit
4852 queries in a single 60s window.

Fix: Set dns.rateLimit.count=0 and dns.rateLimit.interval=0 on both
nodes (private tailnet, no DNS amplification risk). Also persisted in
docker-compose.yml via FTLCONF env vars.

Combined with prior IPv6 fix (4f82520), this resolves all known causes
of DNS flapping.

[#376]
2026-08-06 13:59:45 -05:00
mrcharles bad40feae6 feat(k8s): add worker join script + 6 worker nodes to env config
- env.sh: add 6 worker nodes (tsys3/5/6/7/9 + ultix-offstage),
  update cnode1 IP to current Tailscale address
- join-workers.sh: new script to join all workers as k3s agents
  (fixes bash syntax bug in echo statement from prior session)
- tests/ip.sh: pass-through wrapper for ip command

[#367] [#368]
2026-08-06 13:11:43 -05:00
mrcharles 4f82520e0e fix(netinfra): kill IPv6 on DNS nodes to stop Pi-hole flapping
Root cause of Uptime Kuma DNS up/down alerts: Pi-hole's upstream config
included Google IPv6 DNS (2001:4860:4860::8888), but netinfra-01 has no
IPv6 internet route. Every forwarded query to the IPv6 upstream failed
with "Network unreachable", causing intermittent DNS resolution
failures every ~8 seconds.

Fix applied to both netinfra-01 and netinfra-02:
- Pi-hole upstream set to 8.8.8.8 only (IPv4); removed 192.168.3.16
  (retired netboot) and 2001:4860:4860::8888 (IPv6 Google DNS)
- IPv6 disabled at kernel level (/etc/sysctl.d/99-disable-ipv6.conf)
- knel.net authoritative resolution unchanged (Technitium via revServers)

Verified: zero IPv6 warnings, zero connection errors, DNS resolving
cleanly from all paths after fix.

[#376]
2026-08-06 13:11:38 -05:00
mrcharles d4850a3c02 docs(tailscale): add exit node HA to netinfra-02 for full 01/02 parity
netinfra-02 now advertises exit node routes (0.0.0.0/0, ::/0) in
addition to the subnet route, matching netinfra-01's configuration.
Both nodes are now fully redundant subnet routers + exit nodes.

Subnet route (192.168.0.0/22) is approved and active. Exit node
routes (0.0.0.0/0, ::/0) require separate approval in the Tailscale
admin console.

[#377]
2026-08-06 12:54:45 -05:00
mrcharles 75e229c0ab docs(tailscale): document subnet router HA + fix resolv.conf immutable bug
Adds section documenting the redundant subnet router architecture:
both netinfra-01 and netinfra-02 now advertise 192.168.0.0/22 for
automatic failover. This fixes the SPOF where all remote monitoring
(Uptime Kuma on Cloudron VPS) routed through netinfra-01 alone —
any DERP bounce or Docker LinkChange event took down visibility to
every LAN host simultaneously.

Also documents the resolv.conf immutable flag fix on netinfra-01
(chattr -i) that was preventing tailscaled from updating DNS config
after every link change event.

[#377]
2026-08-06 12:53:36 -05:00
mrcharles 30ddb37b34 fix(netinfra): commit Pi-hole hardening config previously missing from repo
Commit 33b5c76 claimed to harden Pi-hole on both DNS nodes but only
modified markdown — the working docker-compose.yml, gravity-validate.sh,
healthcheck, and autoheal config were never written to the repo, leaving
the DNS hardening unreproducible from version control.

This commits the live, verified-working config from the boxes into
netinfra/pihole/:
- docker-compose.yml (shm_size 1024M root-cause fix, healthcheck, autoheal)
- gravity-validate.sh (pre-start SQLite header check, auto-quarantine corrupt DB)
- .env.example (web UI password templated; real .env gitignored)

Defends against the gravity.db / /dev/shm corruption production outage.
The live password is templated as ${PIHOLE_WEB_PASSWORD} so no secret
enters git.

[#376]
2026-08-06 12:37:46 -05:00
mrcharles 6c58671d15 docs: add Redmine ticket references to subsystem READMEs
Each component README now links to its corresponding Redmine tickets
(closed for completed work, open for pending items) for bidirectional
traceability between code and system of record.
2026-08-06 11:25:19 -05:00
mrcharles 317a98b9b1 docs: migrate work tracking to Redmine system of record
Replace STATUS.md checklists with a Redmine ticket index. All work
status, completion tracking, and issue management now lives in Redmine
(version: Potential to Kinetic Ready, due 2026-09-30). AGENTS.md gains
a Redmine Tracking Policy section with rules for ticket references,
closing procedures, and STATUS.md's new read-only role.
2026-08-06 11:24:48 -05:00
mrcharles 33b5c76bec feat(netinfra): rebuild netinfra-02 as redundant DNS/DHCP + harden Pi-hole
netinfra-02 was accidentally wiped during storage moves. Rebuilt from
netinfra-01 config: Docker + Technitium (167 zones, rsync timer every 60s),
Pi-hole v6 (healthy, gravity.db validated), NTP container, ISC DHCP failover
secondary (load-balancing with primary, 99+84 active leases).

Pi-hole hardened on BOTH nodes with defense-in-depth against the
gravity.db corruption outage:
- gravity-validate.sh: pre-start SQLite header check, auto-removes corrupt DB
- Healthcheck: validates DNS resolution + gravity.db file integrity
- autoheal container: restarts unhealthy containers automatically
- Root cause (/dev/shm too small) already fixed by operator (shm_size: 1024M)

Also corrected DHCP lease count in audit (98 unique active, not 753 raw blocks).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 09:18:45 -05:00
mrcharles d159504121 docs(audit): full ground-truth re-orientation + sync index files to module layout
Read-only audit (nmap/SNMP/SSH/API) reconciling the repo and fleet after the
operator's module/subproject restructure and infrastructure changes. New
authoritative audit doc captures: 4 switches up / 1 down (r2-tor-top offline),
63 VMs (was 28), 167 DNS zones (was 124), 753 DHCP leases (was 333), tsys5
storage now configured, tsys4 upgraded to 9.2.5, k8s cluster DOWN (cnode wiped,
no etcd quorum), netinfra-02 wiped, 3-tier Tailscale fleet (80 nodes).

Index files (README, AGENTS, docmap, STATUS) updated to the new layout
(provisioning→archive, powerman/console/ups→dcinfra, dns-cluster-setup/switches/
dhcp→netinfra, perf→proxmox/perf) and all stale root-level path references fixed.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-08-06 00:49:30 -05:00
mrcharles 04ed433e84 . 2026-08-01 20:31:13 -05:00
mrcharles ea54beea70 meta bits... 2026-08-01 16:45:18 -05:00
mrcharles 959c3a5bb9 more prompt work... 2026-08-01 16:41:13 -05:00
mrcharles 4d86012762 . 2026-08-01 15:50:26 -05:00
mrcharles 6244c1cc25 prep for next ai session 2026-08-01 15:45:23 -05:00
mrcharles 46c35106fb prep for next ai session 2026-08-01 15:44:59 -05:00
mrcharles a1beb6cd3e fix(tuning): persistent NFS tuning fleet-wide + apply-tunings.sh bug fix
Fleet-wide changes (live, documented for audit/AWX):

Persistent tuning via systemd oneshot (all 7 hosts):
- New pfv-nfs-tuning.service runs AFTER tuned.service to override
  16MB TCP buffer caps with 128MB for high-BDP NFS
- New 99-pfv-nfs.conf sysctl fragment (128MB rmem/wmem, 250K backlog)
- Deployed to: tsys1, tsys3, tsys4, tsys5, tsys6, tsys7, tsys9
- Survives reboot — no more tuned clobbering our sysctl values

tsys5 fixes:
- VM 509 stopped, NFS D2/D3/D5 remounted with nconnect=4, VM 509
  restarted (nconnect now live on all tsys5 NFS client mounts)
- SK hynix SSD (sdb1) persisted in fstab at /mnt/ssd2 (UUID-based)
- ssd2 added to Proxmox storage.cfg as dir storage
- Stale S3 export removed from /etc/exports (disk not mounted)
- Broken S4 dir entry removed from storage.cfg
- Duplicate noatime,noatime in fstab fixed
- version=4.2 bug in storage.cfg options fixed (invalid mount option)

tsys4 fixes:
- SSH keys re-standardized to 2 approved (upgrade regenerated host key)
- Tuning persistence working (128MB buffers confirmed live + persistent)

apply-tunings.sh fix:
- Removed version=4.2 from NFS options string (not a valid mount.nfs
  option — causes mount failures; Proxmox auto-negotiates NFSv4.2)

Config files saved to provisioning/ConfigFiles/ for AWX/Ansible consumption:
- 99-pfv-nfs.conf
- pfv-nfs-tuning.service

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-31 08:17:30 -05:00
mrcharles 9c2882ea00 fix: Mermaid diagram syntax + deploy-tuning.sh path + fleet quick fixes
Mermaid diagrams fixed for Gitea rendering:
- Replaced all <br/> → <br> (86 tags across 4 docs)
- Replaced Unicode chars with ASCII (em dash, multiply, middle dot, star)
- Fixed invalid bidirectional dotted arrow in NETWORK-TOPOLOGY.md
  (<--. "..." .--> is not valid Mermaid; replaced with two -.-> arrows)
- Simplified node labels for cleaner rendering

Host fixes applied (live, not in repo — documented here for audit trail):
- SSH keys standardized to exactly 2 approved keys on all 7 hosts
  (removed host-generated RSA keys, added missing VP TechOps key to tsys9)
- Packages installed: iperf3 (tsys9), sysstat (tsys5), nvme-cli (tsys4/5),
  net-tools (tsys1/6/7)
- tsys4 tuning: rmem/wmem 16MB→128MB, netdev_max_backlog 5000→250000,
  tcp_max_syn_backlog 1024→2048, tuned profile→network-throughput
- tsys5: nconnect options added to storage.cfg (3 NFS stanzas patched;
  will apply on next Proxmox NFS remount/reboot)
- noatime applied to root fs on tsys4 + tsys5 (was only tsys5)
- tsys1 lldpd: installed but systemd/dbus issue prevents enabling
  ("Transport endpoint is not connected") — needs investigation/reboot

Bug fix: deploy-tuning.sh computed wrong script path (scripts/ vs
perf/scripts/) — fixed to use dirname/scripts/apply-tunings.sh

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-31 00:12:04 -05:00
mrcharles 290245349f docs(drift): fleet drift report with executive briefing format
Dedicated drift report covering package/service/tuning/security
inconsistencies across all 7 Proxmox hosts. Executive briefing at the
top (what needs decisions, quick severity summary), detailed matrices
in appendixes.

Key drift findings:
- lldpd inactive on tsys1 (blind spot in topology)
- tsys9 missing 2 SSH keys vs fleet standard
- iperf3 missing on tsys9, net-tools missing on tsys1/6/7
- sysstat missing on tsys5, nvme-cli missing on tsys4/5
- tsys4 tuning drift: 16MB TCP buffers, low backlog, wrong tuned profile
- rsyslog + snmpd + beszel inactive fleet-wide (Saturday OAM Day items)
- noatime only on tsys5 root fs (all others use relatime)

Added perf/scripts/probe-drift.sh as a portable reusable drift probe.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 23:27:54 -05:00
mrcharles 0a8865de38 docs: full documentation suite — network, storage, hosts, VMs
Four new standalone reference documents with Mermaid diagrams (renderable
on Gitea), built from live ground-truth data gathered 2026-07-30:

- NETWORK-TOPOLOGY.md: physical topology diagram, complete switch port
  maps (core-sw01 + tor3-stor), VLAN/subnet tables, cross-rack trunk
  analysis with Mermaid, NFS data-flow diagram, per-host NIC reference
- STORAGE-ARCHITECTURE.md: tier definitions, all 14 physical disks
  inventoried with SMART/size/tier, NFS export→mount map with Mermaid,
  per-spindle VM placement (IOPS hotspot analysis), capacity summary
- HOST-REFERENCE.md: one-page-per-host for all 7 nodes with hardware,
  network config, storage, tuning state, role, and running VMs
- VM-INVENTORY.md: VM placement diagram, all 28 running VMs with
  disk/CPU/RAM/config details, optimization audit (cache mode, CPU type,
  qga, disk controller, k8s wnode sizing), stopped VM reference

Updated docmap.md with all four new entries.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 23:10:38 -05:00
mrcharles cd1c2f7332 docs(status): record operator schedule + remote vs onsite work split
Operator onsite Monday 0700. Until then: software/docs/OAM only (Sunday
off). Switch changes + physical cable work deferred to Monday onsite.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 23:03:48 -05:00
mrcharles 24b08aa592 docs(audit): document observability stack + Saturday OAM day plan
Operator-defined monitoring architecture (no tool duplication):
- Beszel (Cloudron VPS): CPU/RAM/disk metrics on all Proxmox hosts via Tailscale
- Uptime Kuma (Cloudron VPS): up/down alerting via Pushover
- LibreNMS (VMID 104, on-prem): SNMP polling for switches (non-Tailsale gear)
- Home Assistant (VMID 100, on-prem): power/temp instrumentation + alerting

Saturday 2026-08-01 is designated OAM Day: goal is 100% instrumentation
coverage across all four tools + centralized syslog.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 23:02:11 -05:00
mrcharles 26de312c42 docs(audit): full ground-truth audit — network, storage, compute, switches
Comprehensive end-to-end audit of the Proxmox cluster covering all three
layers (network, storage, compute) with live data gathered from all 7 hosts
+ 4 switches via SNMP, conman consoles, and LLDP.

Key findings documented in AUDIT-2026-07-30.md:

Network (cross-rack NFS bottleneck root-caused):
- core-sw01 LAG hash = layer-2-3 (no L4 ports) — NFS caps at 1 Gbps per
  host pair regardless of nconnect. tor3-stor already uses layer-2-3-4.
- core-sw01 ch1 member mismatch: g16 is UP but not in LAG; g17 is DOWN.
  Only 3 of 4 cross-rack links active.
- Both switches use static mode (no LACP) for the cross-rack trunk.
- tsys5 bond0 degraded (1 of 2 slaves, nic2 down).

Storage (major topology changes):
- tsys5 new NVMe (CT500P1 500GB) + relocated SK hynix SSD (512GB) — both
  installed but NOT mounted/configured yet.
- tsys4 D3 changed (SK hynix SSD → Seagate HDD, still USB tier-3).
- tsys4 /mnt/albert (Hitachi 2TB internal) + /mnt/backup (WD 5TB USB) added.
- tsys5 S3 unmounted (stale export), S4 gone (broken storage.cfg entry).
- Samsung 860 PRO SSD at 93% full.

Fleet consistency:
- tsys4 behind on PVE 9.1.5/kernel 6.17 (fleet: 9.2.5/7.0.14).
- tsys4/5 bond hash + nconnect + rmem/wmem inconsistent with tsys6/7.
- 10+ VMs using cache=writethrough on NFS (slow synchronous writes).

Gardening: updated STATUS.md, docmap.md, AGENTS.md key scripts table.
Superseded AUDIT-2026-07-29 as the authoritative audit.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 23:01:43 -05:00
mrcharles d9b32e4eef fix(console): fix udev symlink naming bug + add portable audit tooling
Console fix: generate-config.sh wrote SYMLINK+="console/$name" (singular)
but ser2net.yaml opens /dev/consoles/$name (plural). They never matched,
so after every reboot the console ports failed until setup.sh's manual
fallback re-created the symlinks. Fixed the udev rule to use "consoles/"
to match ser2net and the README.

New portable read-only audit tools (AGPLv3-friendly, config-driven):
- perf/scripts/probe-storage.sh: disk/mount/export/SMART/storage.cfg probe
- perf/scripts/probe-network.sh: NIC/bond/LLDP/NFS/nconnect probe
- perf/scripts/conman-console.py: PTY-based conman console driver (replaces
  the old sw-capture.py that conflicted with ser2net)
- perf/scripts/snmp-switch-audit.py: SNMP-based switch inventory (interfaces,
  LLDP, LAG, VLANs) via pysnmp or net-snmp

Removed stale pre-conman switch tooling (sw-capture-remote.sh, sw-capture.py,
sw-probe.sh, sw-conman-probe.sh) and old .cmds files. Added fresh .cmds
files for the two cross-rack trunk endpoint switches.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 22:51:27 -05:00
mrcharles e7312ad460 chore: retire pfv-tsys8 and fix stale host-skip comment in deploy-check.sh
pfv-tsys8 is permanently retired. Strike its only doc reference and
correct the stale comment in perf/deploy-check.sh that wrongly listed
pfv-tsys9 as "off the air" (tsys9 is online and in the host array;
tsys8 is the retired host).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 22:30:24 -05:00
mrcharles 48cdd81779 docs(ups): document apcupsd test results and Energy Dashboard limitations
The APC Smart-UPS C 1500 (FW 02.2, mfg 2012) does not expose ups.load or
ups.realpower over USB HID. Tested apcupsd as an alternative — it reads the
same HID descriptor and returned COMMLOST. Debian's apcupsd package also
conflicts with nut-server (mutually exclusive).

Documented that the HA NUT integration provides battery/runtime/status sensors
only, not wattage for the Energy Dashboard. Practical path is a smart plug on
the UPS output. Updated daily-ops examples (removed ups.load reference since
the variable doesn't exist on this UPS).

NUT stack fully restored after the apcupsd test cycle.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 12:02:15 -05:00
mrcharles 4b26aca5ee feat(ups): add Home Assistant NUT integration via REST config-flow API
Drive HA's REST config-flow endpoint to add the NUT integration programmatically,
no manual UI clicks required. The script (setup-ha-nut.sh + ha-nut-setup.py) is
idempotent — skips if a NUT entry already exists.

Key finding: HAOS runs Tailscale as an isolated add-on container, so the HA core
container cannot route to Tailscale IPs. Added a LAN listener (192.168.3.11:3493)
to upsd so HA can reach it over the shared vmbr0 bridge. Both VMs (pfv-bms HA at
192.168.3.12 and pfv-tsys1 at 192.168.3.11) are on the same bridge.

Integration is live — sensors for battery charge (100%), status (Online), and
status data (OL) are reporting.

Secrets (HA token, NUT password) are read from ~/.config/pfvcluster/ and never
committed to the repo.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 11:42:25 -05:00
mrcharles 5456c783c6 feat(ups): NUT setup for APC Smart-UPS C 1500 on pfv-tsys1
Deploy Network UPS Tools to monitor the server-room UPS units and feed data to
Home Assistant for power-usage tracking.

APC Smart-UPS C 1500 is fully operational: battery charge, runtime, load,
voltage, and status all reporting via usbhid-ups + APC HID 0.100 subdriver.
upsd listens on Tailscale (100.121.189.98:3493) for HA polling; local upsmon
provides graceful hypervisor shutdown on battery-low.

Key deployment fix: the udev rule must cover SUBSYSTEM=="usb" (raw
/dev/bus/usb) in addition to hidraw, because usbhid-ups opens the raw USB
device file after dropping to the nut user via setuid(). Matching only hidraw
causes EACCES on driver start.

Tripp Lite UPS (09ae:3016) is blocked — driver matches TrippLite HID 0.85
subdriver but fails reading the 878-byte HID Report Descriptor (EAGAIN via USB
hub). Needs physical reseat to direct motherboard USB port. Driver masked to
prevent restart-loop spam.

Files: ups/discover.sh, ups/setup.sh (idempotent), ups/status.sh, ups/README.md

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 09:34:05 -05:00
mrcharles 0fa0692c37 chore: enforce shellcheck across the repo
Establish shellcheck as a mandatory pre-commit quality gate and bring all 93
shell scripts to a clean state.

- tests/shellcheck.sh: wrapper that runs koalaman/shellcheck:stable via Docker
  (no native binary needed), skips vendored + upstream librenms-agent scripts.
- .shellcheckrc: documents intentional codebase-wide disables (dynamic source
  paths SC1090/SC1091, client-side ssh expansion SC2029).
- AGENTS.md: new Git Policy rule mandating clean shellcheck for every shell
  script before commit.

Fixes applied (real bugs + quality): missing quote in netinfra/gather-configs.sh
(caused cascading parse errors), unquoted expansions, declare-and-assign masking,
egrep -> grep -E, $FUNCNAME array indexing, unused variable removal, cd || exit.
Intentional patterns (sourced config, sysfs/ps diagnostics, ssh heredocs that
expand local config) get justified targeted disables.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 08:56:31 -05:00
mrcharles 54e9927167 docs(status): mark console validation complete with pfv-r2-sw known issue
6/7 consoles validated working via conman from workstation. pfv-r2-sw
(port 2007) produces zero bytes at all baud rates — adapter is present
and correctly mapped but the device is physically unreachable (powered
off, cable disconnected, or console port disabled). Requires onsite
investigation.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-30 08:31:46 -05:00
mrcharles 72ab6b4b32 docs(status): add UPS USB tracing action item
APC + TrippLite UPS units not enumerating on tsys1/tsys4. tsys4 shows no HID
device at all; tsys1 may have USB passed through to pfv-bms. Need to trace
cables and confirm physical connection before NUT driver setup.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 19:53:34 -05:00
mrcharles 25a7234f40 fix(console): switch ser2net accepter to telnet(rfc2217) for proper NVT
The original ser2net config used raw TCP accepters, but conman's
dev="host:port" speaks telnet protocol. This protocol mismatch caused
conman's telnet NVT state machine to strip bare CR characters from
device output — particularly from Dell switches that send \n\r (LF+CR)
line endings — producing stair-stepped and garbled terminal display.

Fix: change ser2net accepter from `tcp,IP,PORT` to
`telnet(rfc2217),tcp,IP,PORT`. With both sides speaking telnet, binary
mode is negotiated and CR/LF translation is handled correctly by the
telnet NVT layer. RFC2217 also enables baud rate negotiation for
future use.

Validated at the byte level: PTY capture confirms all prompts render
cleanly at column 0 with proper CRLF handling across 5+ Enter presses.
All 7 consoles connected, 7 log files actively capturing.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 19:35:05 -05:00
mrcharles 51db52f862 security: redact plaintext password from all files
Replace all occurrences of the fleet admin password with REDACTED_PASSWORD.
Git history will be rewritten next to purge it from all prior commits.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 19:28:22 -05:00
mrcharles 050e235f8b feat(awx): deploy AWX 24.6.1 on tsys-awx via k3s + AWX Operator
AWX (Ansible automation controller) deployed on tsys-awx.knel.net (VMID 600
on pfv-tsys6). Accessible at http://tsys-awx.knel.net:80 (admin/REDACTED_PASSWORD).

Deployment details:
- k3s v1.36.2 single-node (no Traefik) on Debian 13 trixie
- AWX Operator 2.19.1 + AWX CR (LoadBalancer service via k3s ServiceLB)
- PostgreSQL 15 bundled (8 GiB PVC on local-path provisioner)
- VM disk resized 32->60 GB; swap partition converted to swapfile
- Fixed kube-rbac-proxy image (gcr.io removed -> quay.io/brancz replacement)
- Fixed AWX CR field names (postgres_data_volume_size ->
  postgres_storage_requirements)

Scripts: awx/install-k3s.sh, awx/deploy-awx.sh, awx/verify-awx.sh
Manifests: awx/awx-instance.yaml, awx/namespace.yaml
Docs: awx/README.md

All pods running: operator 2/2, postgres 1/1, web 3/3, task 4/4.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 19:12:03 -05:00
mrcharles fa0c77fb58 fix(console): switch conman from ser2net-telnet to direct serial access
The ser2net layer between conman and the serial devices was causing
stair-stepping in terminal output. conman's dev="host:port" uses telnet
protocol, but ser2net's accepter was raw TCP (no telnet mode). The telnet
NVT state machine in conman was stripping bare CR characters from device
output — particularly from Dell switches that use old-style \n\r (LF+CR)
line endings instead of standard \r\n. In telnet, a bare \r not followed
by \n or NUL is non-compliant and gets dropped, leaving bare \n that
causes stair-stepping in the raw-mode terminal.

Fix: conman now opens serial devices directly via the stable udev
symlinks (/dev/consoles/<name>) with seropts, eliminating the telnet
layer entirely. ser2net is stopped and disabled but remains installed
for emergency TCP access (documented workflow: stop conmand, start
ser2net, use telnet, then reverse).

Architecture change:
  Before: device → serial → ser2net (raw TCP) → conman (telnet NVT) → terminal
  After:  device → serial → conman (direct) → terminal

Verified: 7/7 serial devices held by conmand, 7/7 log files capturing,
all \r bytes preserved in terminal output (confirmed via PTY capture).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 18:29:31 -05:00
mrcharles 80693e4f8e feat(awx): scaffold AWX deployment scripts and manifests
Create top-level awx/ directory with k3s install script, AWX operator +
instance deployment script, namespace manifest, and AWX custom resource
(LoadBalancer service type, bundled PostgreSQL on local-path).

Target: tsys-awx.knel.net (VMID 600 on pfv-tsys6, 4c/12GB/60GB disk).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 17:45:26 -05:00
mrcharles 396349640f feat(netinfra): migrate DHCP from pfv-netboot to netinfra-01/02 with failover
ISC DHCP server with failover-peer (primary/secondary) deployed on
pfv-netinfra-01 (192.168.3.252) and pfv-netinfra-02 (192.168.3.253). Both
servers in "normal" state, load-balancing 333 active leases.

Migration details:
- Copied all 37 host reservations + subnet/pool config from pfv-netboot
- DHCP lease database copied from netboot for zero-disruption cutover
- DNS servers changed from 192.168.3.250 (netboot) to 252/253 (netinfra pair)
- NTP servers (252/253) added to DHCP options (netboot didn't hand out NTP)
- Netmask on both nodes fixed /24 -> /22 to match the network
- Webmin + DHCP module installed on both nodes (port 10000, SSL)
- pfv-netboot DHCP stopped + disabled
- Tested via sectestbed-sandbox (DHCP lease obtained from 252, verified DNS/NTP/gateway)
- Snapshot "pre-dhcp-migration" on sandbox as rollback point

Configs: netinfra/dhcp/dhcpd-{primary,secondary}.conf
Plan + results: netinfra/dhcp-migration.md

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 17:22:15 -05:00
mrcharles 9c47d93a2f docs(goals): codify local storage policy — ephemeral scratch, not OS drives
Add CAPACITY-GOALS.md §6: local storage is a scratch tier for ephemeral workloads
(container image cache, build artifacts, emptyDir), NOT for VM OS drives (which
stay on NFS where redundancy is handled via 01/02 pairs + cross-disk spreading).
Documents per-host local storage suitability: tsys3/9/2 excellent (NVMe/SSD),
tsys5 usable (HDD bulk), tsys6/7 unusable (USB 2.0 at 30 MB/s). Includes the
future k8s StorageClass mapping (local-fast for scratch, nfs-hdd for persistent).
Adds forward-looking §12 to the audit referencing this policy.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 16:33:56 -05:00
mrcharles edb34a75ec docs(status): sync action items with audit, fix stale disk refs
Rewrite STATUS Pending section to match the audit's phased action plan (§11)
exactly. Remove stale "wnode-tsys6 D5->S4, wnode-tsys7 D5->S1" line (wnode6
already on S4, wnode7 on S3, S1 is not a target). Fix section reference (§12 ->
§11). Correct actual-usage percentages (73%/37%/62%) and CPU idle range (0-5%).
Add D2 IOPS hotspot as a known issue. Fix sectestbed suite count in audit (x9 ->
x10). Verified all VM counts and host data against fresh ground truth.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 16:29:17 -05:00
mrcharles 23e5ecc4cb docs(audit): mark tsys6/tsys9 S1 storage.cfg fixes as resolved
Operator corrected the tsys6 S1 export path (S21 -> S1) and added the missing
S1 entry to tsys9. Verified both via pvesm status: all 8 NFS stores active on
each host. Updated audit findings table, §5.1/§5.2, and Phase 0 checklist.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 16:23:37 -05:00
mrcharles 915111e9dd docs(audit): consolidate to single audit with corrected storage topology
Delete AUDIT-2026-07-28.md (old audit). Rewrite AUDIT-2026-07-29.md as the
sole authoritative audit after reading live /etc/pve/storage.cfg, /etc/exports,
and /etc/fstab on every host.

Key corrections:
- local-nonprod on tsys5 IS the S1 physical disk (dir storage, not separate).
  S1 carries 16 sectestbed VMs, not 0 — it is NOT a migration target.
- storage.cfg is per-host, not cluster-wide — each node defines its own entries.
- BUG: tsys6 S1 export path typo (S21 instead of S1) — S1 broken on tsys6.
- GAP: tsys9 has no S1 defined in storage.cfg.
- Revised disk migration plan avoids S1, targets S2/S3/S4/D5 for IOPS spreading.
- D2 goes 17->5 VMs; S2/S3 absorb the load (7 each); S1 untouched (16 idle VMs).
- Refreshed ground truth: cnode5 deleted, cnode4 stopped (user actions this session).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 14:32:43 -05:00
mrcharles 6199368710 docs(audit): add VM placement plan (current/target inventory + action checklist)
Add §13-§15 to AUDIT-2026-07-29: current per-VM host+disk inventory, target
end-state inventory (RAM/CPU relief + HDD IOPS rebalancing), and a phased action
checklist. Headline disk finding: D2 carries 17 of 26 NFS VMs while S1 sits
empty — the plan rebalances to 4-5 VMs per spindle (D2 17->5, S1 0->5) using HDD
exports only. Host moves relieve tsys1 (115%->77% alloc) and tsys9. All disks
stay on HDD (SSD/NVMe tiering deferred to post-Friday).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 06:57:51 -05:00
mrcharles 575d79f611 docs(audit): RAM/CPU placement audit + codify capacity goals
Add CAPACITY-GOALS.md (80% RAM / ~50% idle CPU targets + workload placement
model) and AUDIT-2026-07-29.md (live end-to-end analysis of all 7 hosts). Key
findings: memory goal met fleet-wide (highest 72%); CPU idles 0-6% (chase via
workloads, not hardware); tsys5 is single-CPU not dual; cnode4/5 are stale
running VMs. Answers the wnode-sizing question (grow tsys9, not tsys1), confirms
tsys6/7 for RackRental, and gives a DDR3 compatibility decision tree (RDIMM-only
servers reject consumer UDIMM). Refreshed STATUS/docmap and the returned audit
logs, and fixed an empty-array abort in deploy-check.sh.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 05:32:12 -05:00
mrcharles 428e9ecb1d refactor(docs): co-locate k8s/proxmox docs with their code + fix broken perf paths
Docs moved from docs/k8s -> k8s/docs and docs/proxmox -> proxmox/docs to sit
beside their code. Updated all cross-references (docmap, STATUS, README, AGENTS,
k8s README). Fixed SCRIPT_DIR path resolution in 3 perf scripts broken by the
reorg (deploy-check/and-fix/finish-host pointed at root/scripts instead of
perf/scripts). Registered the new top-level proxmox/ dir in all listings.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-29 05:28:07 -05:00
mrcharles b860cb116d docs(audit): revised fleet audit with corrected placements + Friday plan
User has fixed the HA pair storage redundancy issue:
- netinfra-01 moved to S2 (tsys5), netinfra-02 stays on D2 (tsys4)
- ucs-02 moved to S2 (tsys5), ucs-01 stays on D2 (tsys4)
Both pairs now survive any single storage server failure. Marked as
FIXED in Known Issues.

Revised audit includes:
- Corrected VM placements from fresh qm config data
- Host capacity analysis: tsys1 (118% RAM), tsys5 (135% RAM), tsys9
  (131% RAM) are overcommitted; tsys6 (38%) and tsys7 (30%) have room
- Wnode storage rebalancing plan: move wnode-tsys6 → S4, wnode-tsys7 →
  S1, giving every wnode its own physical disk
- Friday maintenance plan: tsys2 Proxmox install as primary deliverable,
  VM migrations to relieve overloaded hosts, USB-passthrough VMs relocate
  from tsys1 to tsys9
- Cnode quorum: documented as fundamental constraint (3 cnodes, 2 NFS
  servers), accepted for now, revisit post-Friday when local-SSD boot
  becomes available

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 21:09:56 -05:00
mrcharles b6f94483e5 fix(console): enable conmand remote access + add conman client script
conmand was binding to localhost only (server loopback=on default), so
the conman client on workstations couldn't connect. The intended workflow
is: conman client (workstation) → conmand (pfv-tsys4:7890 over Tailscale)
→ ser2net (TCP 2001-2007) → serial device. Without remote conmand access,
users had to telnet directly to ser2net, which conflicts with conmand's
persistent connections (kickolduser kicks the telnet session immediately).

Changes:
- generate-config.sh: add server loopback=off to conman.conf so conmand
  listens on 0.0.0.0:7890 (reachable via Tailscale)
- query-remote.sh: new script for workstations — installs conman client,
  verifies connectivity, lists or connects to consoles
- README.md: clarify access model (conman primary, telnet emergency only
  with conmand stopped). Document the kickolduser conflict.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 20:07:17 -05:00
mrcharles 28e0b0c7a6 docs(audit): fresh fleet audit + fix stale paths across 13 perf scripts
Fresh Proxmox fleet audit (2026-07-28) with current VM placements, RAM,
CPU, and storage for all 7 reachable hosts. Written to
docs/proxmox/AUDIT-2026-07-28.md — supersedes placement data in
PROJECT.md sections 4-8.

Key audit findings:
- CRITICAL: UCS01/02 and netinfra01/02 HA pairs both still on tsys4
  storage. tsys4 failure = DNS/DHCP/NTP + LDAP/AD fully dark. These
  migrations were the #1 recommendation from the previous audit and
  have not been done.
- CRITICAL: 2 of 3 active k3s cnodes (cnode1 + cnode2) on tsys4 NFS.
  tsys4 failure = etcd quorum lost.
- 59% of running VMs still on tsys4 storage (improved from 68%).
- cnode VMIDs have changed since PROJECT.md was written (cnode1 is now
  VMID 906 on tsys9, cnode2 is VMID 705 on tsys7, etc.)

Gardening fixes:
- Removed duplicate fleet-audit.sh (check.sh + deploy-check.sh already
  exist for this purpose)
- Fixed hardcoded path /home/reachableceo/projects/perfopt in 13 perf/
  scripts to use BASH_SOURCE-derived relative paths (per AGENTS.md
  self-locating scripts convention)
- Updated STATUS.md Known Issues with the two critical findings
- Updated STATUS.md Pending with prioritized pre-k8s action items
- Registered AUDIT-2026-07-28.md in docmap.md

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 20:07:09 -05:00
mrcharles 0600321968 fix(powerman): use -h flag for remote server in query script
Debian's powerman client uses the -h/--server-host flag, not the
POWERMAN_SERVER env var, to connect to a remote daemon. Update query-remote.sh
to pass -h explicitly.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 19:48:02 -05:00
mrcharles 5dcddfa4ba feat(powerman): add remote query script for workstation PDU access
Script for any Tailscale-connected workstation to install the powerman
client and query the Cyclades PDU on pfv-tsys1. Sets POWERMAN_SERVER so
all powerman commands route to the remote daemon. Handles missing sudo
gracefully with instructions for manual install.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 19:47:01 -05:00
mrcharles 322e20e50f fix(powerman): bind to Tailscale + localhost instead of 0.0.0.0
Change powermand listen address from 0.0.0.0:10101 (all interfaces) to
127.0.0.1:10101 (local admin) + Tailscale IP:10101 (remote access). The
setup.sh now auto-detects the Tailscale IP at deploy time.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 19:45:17 -05:00
mrcharles eb99b00a70 feat(console): manage 7 switch consoles via ser2net+conman on pfv-tsys4
Solve the long-standing USB adapter enumeration shift problem: 9 Prolific
USB-to-DB9 adapters on pfv-tsys4 have no unique serial numbers and get
assigned /dev/ttyUSB0-8 based on enumeration order, which changes on every
reboot and breaks the old /root/conmap + manual screen workflow.

Solution: udev rules pin each adapter by its ID_PATH (physical USB port
topology), which is stable across reboots regardless of enumeration order.
Each adapter gets a named symlink in /dev/consoles/<name>. ser2net opens
these stable symlinks and exposes them on TCP ports (2001-2007) bound to
the Tailscale interface only. conman connects to those TCP ports for
session logging and multi-user console sharing.

Architecture (layered, no port sharing):
  USB adapter → udev symlink → ser2net (TCP) → conman (logging + mux)

Port assignments (all on Tailscale IP 100.70.77.93):
  2001 = pfv-core-sw01     2002 = pfv-tor3-mgmt    2003 = pfv-tor3-stor
  2004 = pfv-rrinfra-rtr   2005 = pfv-r2-tor-top   2006 = subodev-torsw
  2007 = pfv-r2-sw

Scripts (console/):
- mapping.txt: source of truth (TCP port | name | ID_PATH | baud | comment)
- generate-config.sh: generates udev rules, ser2net.yaml, conman.conf
  entries from mapping.txt. Idempotent (markers in conman.conf for clean
  regeneration). Uses | delimiter (ID_PATH values contain colons).
- setup.sh: full deploy — generate configs, create symlinks (udev trigger
  + manual fallback for already-discovered devices), create conmand
  systemd unit (Debian doesn't ship one), restart services
- discover.sh: read-only USB adapter and service state discovery
- validate-conman.sh: verify conman→ser2net→device data path and log capture

Issues fixed during development:
- /dev/console is a kernel char device (major 5, minor 1) — cannot create
  a directory there. Changed symlink namespace to /dev/consoles/.
- conman 0.3.x has no 'include' directive — CONSOLE entries written
  directly into /etc/conman.conf between idempotent markers.
- Debian conman package has no systemd unit — created
  /etc/systemd/system/conmand.service with After=ser2net ordering.
- conman.conf had no LOGDIR — logs weren't being written to
  /var/log/conman/. Fixed by adding server logdir directive.

Validation: 7 symlinks resolving, 7 TCP ports on Tailscale, conmand with
7 consoles registered, 7 log files actively capturing console output,
both services enabled for reboot survival.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 19:44:02 -05:00
mrcharles 799d270261 docs: fix gardening variance + strengthen protocol for new directories
Two variances from the gardening protocol were found during self-audit:
the k8s/ and powerman/ directories were added but the root README.md
directory table and the AGENTS.md Key Scripts table were not updated.

Root cause: the gardening protocol enumerated STATUS.md and docmap.md but
did not explicitly call out the root README.md directory table or the
AGENTS.md Key Scripts table, so they were easy to miss when adding a new
top-level directory.

Fixes:
- README.md: add k8s/ and powerman/ to the Directory Structure table
- AGENTS.md: add install-cp.sh and powerman/setup.sh to Key Scripts

Permanent fix (encode in protocol so it cannot recur):
- Add rule 5 to the Automatic Gardening Protocol: when a new top-level
  directory is created, ALL directory listings must be updated
  (README.md table, AGENTS.md layout block, AGENTS.md Key Scripts)
- Add rule 6: a grep-based self-audit command to run before commit,
  verifying the new directory appears in all four canonical files

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 18:49:39 -05:00
mrcharles 8124483da8 feat(powerman): manage Cyclades PM10i PDU via powerman on pfv-tsys1
Set up centralized PDU management for a Cyclades AlterPath PM10i (10
controllable AC outlets) connected to pfv-tsys1 via a Prolific USB-to-DB9
serial adapter. powermand is now listening on 0.0.0.0:10101, making the
PDU manageable over the network from any host on the tailnet.

Scripts (powerman/):
- discover.sh: gather USB adapter, powerman state, device definitions
- setup.sh: idempotent setup — udev rule (stable symlink by serial number),
  powerman.conf with 10 outlet nodes, fix powermand dialout group, restart
  service. Overridable via env vars for other hosts/PDU types
- test-pdu.sh: validate control by cycling outlet 10 off then on (8/8 pass)
- status.sh: quick PDU status check

Issues fixed during setup:
- Config pointed at /dev/ttyUSB0 but adapter is at /dev/ttyUSB1 (fixed
  with udev symlink /dev/cyclades-pm10 pinned to adapter serial)
- powermand (user:powerman) lacked dialout group membership to open the
  serial device (fixed with usermod + udev GROUP="dialout")

Validation: outlet 10 turned off (confirmed), turned on (confirmed), then
cycled. All 10 outlets currently ON and manageable.

TODO tracked for Friday: rename outlets from generic (outlet-1..10) to
match physical devices, and change PDU admin password from factory default.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 18:22:47 -05:00
mrcharles 8f1642bf96 feat(k8s): deploy 3-node k3s HA control plane over Tailscale
Bootstrap a regular (non-ITAR) k3s cluster on cnode1/2/3 with embedded
etcd. All cluster communication — node registration, API server, etcd
peering, flannel VXLAN — runs exclusively over Tailscale IPs. Zero LAN
addresses appear in node status or TLS certificates.

Scripts (k8s/):
- env.sh: shared config (Tailscale IPs, SSH opts, k3s version)
- wipe.sh: remove existing k3s from all cnodes
- install-cp.sh: full bootstrap (cnode1 --cluster-init, then cnode2/3 join)
- join-servers.sh: re-join cnode2/3 only (fixes broken join state)
- post-setup.sh: apply NoSchedule taints, fetch kubeconfig, verify
- verify.sh: 13-point health check (nodes, Tailscale IPs, taints, etcd,
  CoreDNS, API server, workload isolation)
- probe-nodes.sh: SSH + Tailscale reachability check

All 3 cnodes are tainted control-plane:NoSchedule so no user workloads
can schedule on the control plane. 13/13 health checks pass.

Docs updated: k8s README TL;DR reflects k3s (not Talos) as the deployed
choice, with Talos preserved for the future ITAR cluster.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 12:21:33 -05:00
mrcharles 174deea2a9 docs(k8s): add Talos architecture, distro decision, and bootstrap plan
Author the docs/k8s/ directory capturing the pfv-k8s control-plane design:

- README.md: TL;DR of all decisions (distro, runtime, cnode count, admin
  access, identity, tenancy, registry, storage)
- DISTRO-DECISION.md: Talos vs k3s analysis. Recommend Talos because the
  ITAR/classified requirement makes its immutable, API-only, measured-boot
  posture structurally easier to certify than SCAP-hardened Debian. k3s was
  only ever a plan (no cluster deployed), so cutover cost is ~zero.
- ARCHITECTURE.md: target arch with mermaid diagrams covering control
  plane, LAN-only network with Tailscale subnet-router admin, Cilium CNI,
  OIDC to Keycloak, per-tenant vcluster isolation (incl. ITAR tenant),
  Harbor pull-through cache on D3 SSD, bootstrap sequence, and DR.

Gardening: register docs/k8s/ in docmap.md, update STATUS.md with the new
k8s section and the three pending user decisions (cnode count, host spread,
Tailscale pattern).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 11:50:56 -05:00
mrcharles 99270c344d docs(agents): enforce always-commit-push policy
Strengthen Git Policy point 1 so agents never hold work for review. The
user reviews rendered markdown on Gitea after push, so pausing to "let
them read first" defeats the workflow. Explicitly overrides any default
conservative commit-and-hold behavior.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 11:50:51 -05:00
mrcharles 6b6a8ff021 docs: organize docs into project subdirectories
Split docs/ into project-based subdirectories:
- docs/proxmox/     fleet ops, hardware, k8s (PROJECT.md, TODO.md, K8S.md)
- docs/server-build/ provisioning, security, DNS (SECURITY.md, tailscale.md,
                    DEPLOYMENT.md, TSYS-2FA-GUIDE.md, DEVELOPMENT-GUIDELINES.md)
- docs/archive/     historical AI reviews, completed todos, pre-refactor docs

docmap.md rewritten with new paths. All cross-references in AGENTS.md,
README.md, STATUS.md, and dns-cluster-setup/README.md updated.

Code directories unchanged — scripts stay where BASH_SOURCE expects them.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 11:34:23 -05:00
mrcharles 4f2bfe18a5 docs: add STATUS.md, docmap.md, encode gardening protocol
Restructure top-level to exactly three .md files:
- AGENTS.md (agent operating instructions)
- README.md (project overview, links to status + docmap)
- STATUS.md (living project status, agent-maintained, human read-only)

Add docs/docmap.md as the single documentation index/map. All docs are
categorized (active, operational guides, historical) with last-reviewed
dates. Includes the agent gardening protocol requiring agents to update
STATUS.md and docmap.md after every work session.

Rewrite AGENTS.md to be lean: points to docmap.md for doc discovery,
encodes the automatic gardening protocol (keep docs/code in sync, grep
for stale paths after renames, update STATUS.md after infrastructure
changes). All references are Gitea-renderable relative links.

Simplify README.md: header links to STATUS.md + docmap.md + AGENTS.md,
doc table replaced with pointer to docmap.md.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 11:32:03 -05:00
mrcharles a7fbad60ab docs: end-to-end gardening — links, stale refs, tailscale.md update
Comprehensive documentation gardening across the merged repo:

- tailscale.md: fully rewritten with current ground truth. The netinfra
  pair now runs production Technitium with all knel.net records
  replicated. Both LAN IPs resolve knel.net device names and recurse
  externally. The old "NXDOMAIN / zone is stale" findings are replaced
  with the resolved state and current recommendations.
- AGENTS.md: rewritten with Gitea-compatible clickable relative links
  to all key scripts and docs. Autonomous commit/push policy
  prominently documented. SSH user corrected to localuser.
- README.md: directory table and docs table now use clickable links.
- All .md cross-references converted to Gitea-renderable relative links.
- Stale path references (ProjectCode/, Project-Tests/, ProjectDocs/)
  updated to current names (provisioning/, tests/) across all docs.
- Stale repo name "FetchApply" / "KNELServerBuild" updated to
  "PFVCluster" in actionable docs; historical AI-review docs tagged
  with an HTML comment notice.
- REFACTORING-EXAMPLES.md: tagged as historical (pre-refactor patterns).
- tests/README.md, dns-cluster-setup/README.md, docs/DEPLOYMENT.md,
  docs/SECURITY.md: path references fixed to current structure.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 11:28:52 -05:00
mrcharles c2b592a66b docs: unified README and AGENTS.md for merged repo
Replace the KNELServerBuild README with a unified PFVCluster README
covering both provisioning and cluster ops. Update AGENTS.md to document
the merged repo layout, key scripts, and project context. Consolidate
all documentation under docs/.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 11:25:16 -05:00
mrcharles 4851517947 refactor: reorganize merged repo into clean directory structure
Reorganize the merged KNELServerBuild + PFVCluster repo:

  provisioning/    server provisioning (was ProjectCode/ +
                   Project-Includes/ + Project-ConfigFiles/)
  tests/           test suite (was Project-Tests/)
  perf/            Proxmox perf scripts (was top-level *.sh + scripts/)
  docs/            all documentation (was ProjectDocs/ + PROJECT.md +
                   K8S.md + TODO.md)
  dns-cluster-setup/  Technitium DNS cluster (unchanged)
  netinfra/        netinfra audit scripts (unchanged)
  switches/        switch configs (unchanged)
  vendor/          vendored KNELShellFramework (unchanged)

Update all internal path references from old directory names
(ProjectCode/, Project-Includes/, Project-Tests/) to the new ones
(provisioning/, tests/) across all scripts.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 11:24:39 -05:00
mrcharles 132c0854d1 merge: combine KNELServerBuild into PFVCluster
Merge the KNELServerBuild repository (server provisioning, security
hardening, DNS cluster setup, test suite) into PFVCluster (Proxmox
cluster ops, performance tuning, fleet audit). Both histories are
preserved via --allow-unrelated-histories.

The two repos had no source-file collisions; only AGENTS.md and
.gitignore conflicted (both resolved by merging content from both).

Directory reorganization and doc gardening will follow in subsequent
commits.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 11:23:22 -05:00
mrcharles b1088e8487 feat(dns-cluster): replicate Technitium production to netinfra pair
Set up a fully scripted, documented Technitium DNS cluster that
replicates the production instance from tailscale-router to
pfv-netinfra-01 (primary) and pfv-netinfra-02 (secondary).

What it does:
- EXPORT: reads the production Technitium config (auth.config with
  users + 2FA, dns.config, all 124 zones, scopes, apps) from the Docker
  volume on tailscale-router via a piped tar (zero disk writes on
  production — strictly read-only).
- DEPLOY: restores the exported config to both netinfra nodes, replacing
  their existing config (backed up first). Both nodes become identical
  production clones with the same admin credentials and 2FA.
- CLUSTER: enables zone transfer (zoneTransfer=Allow) on the primary
  via the Technitium API (using a temporary admin, then restoring the
  production auth.config). Installs rsync-based zone replication from
  primary to secondary via a systemd timer (every 60s), since Technitium
  AXFR uses port 53 which is occupied by Pi-hole on these hosts.
- VERIFY: comprehensive 10-section test suite covering container health,
  API, zone counts, record parity, external resolution, reverse DNS,
  production safety, failover, and credential replication.

Scripts:
- remote-dns.sh: SSH chokepoint for all DNS host access
- setup.sh: master orchestrator (export → deploy → cluster → verify)
- sync-zones.sh: rsync-based zone replication (installed as systemd timer)
- verify.sh: 10-section verification suite

Safety:
- tailscale-router is NEVER modified (read-only export only)
- Production auth.config is backed up before any temporary admin swap
- Each node's existing config is backed up before replacement
- The export tarball is gitignored (contains production credentials)

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 08:50:14 -05:00
mrcharles 1951667f8b fix(network): remove interface restriction that broke NTP client sync
The ntp.conf hardening used `interface ignore wildcard` +
`interface listen 127.0.0.1`, which binds ntpd to loopback only. Outbound
NTP queries to the upstream servers then carried a 127.0.0.1 source
address that the servers cannot reply to, so the daemon's peers stayed
stuck in .INIT. with reach 0 — even though the servers are reachable
(verified: ntpdate -q succeeds, ntpd does not).

Replace the interface-based restriction with restrict-based hardening:
`restrict default ignore` blocks unsolicited queries from any host (so
the box never serves time to others), while explicit allow rules for the
two upstream servers and localhost let the client sync normally.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 06:03:21 -05:00
mrcharles 9a4961d94b docs(network): analyze Tailscale vs managed DNS conflict
Add an architecture analysis for the tension between Tailscale's
default resolv.conf management (100.100.100.100) and the managed
LAN-resolver resolv.conf (.252/.253). Documents a key finding from
live-network probing: knel.net device records only resolve via the
Tailscale MagicDNS path; querying the LAN DNS servers directly returns
NXDOMAIN because their knel.net zone is stale (SOA serial 2025-06-23).

Lays out four options (Tailscale-owned, LAN-pinned, split DNS,
Tailscale-pushes-LAN-resolvers) with pros/cons, recommends leaving DNS
to Tailscale in the short term (since wazuh/postfix/syslog depend on
knel.net names that only resolve there) and fixing the Technitium/Pi-hole
knel.net zone before pinning the LAN resolvers. Confirms the NTP
(LAN-IP) change is safe regardless. Flags that the managed-resolv.conf
change will be overwritten by Tailscale and would break knel.net
resolution if it ever sticks.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 05:31:47 -05:00
mrcharles f010fa9609 feat(network): use pfv-netinfra-01/02 as redundant DNS and NTP
Route every host built by this project through the new
pfv-netinfra-01 (192.168.3.252) / pfv-netinfra-02 (192.168.3.253)
pair for both name resolution and time, with automatic failover.

- NTP: replace the single pfv-netboot.knel.net upstream with both
  netinfra servers (iburst) so time sync survives either one failing.
- DNS: add a managed static /etc/resolv.conf (new ConfigFiles/Resolv/).
  The repo previously had no resolver configuration at all. Both servers
  are listed so glibc falls through to the secondary on failure.
- DHCP: request domain-name-servers/domain-search/ntp-servers and
  supersede them to the netinfra pair, so a DHCP renew can't silently
  revert to whatever the DHCP server advertises.
- SetupNewSystem.sh: deploy resolv.conf (robustly replacing any
  systemd-resolved/NetworkManager symlink) and add pfv-netinfra to the
  NTP-server self-exclusion guard so those boxes don't client off
  themselves.

LAN IPs are used throughout (not the knel.net hostnames) because those
hostnames resolve to Tailscale CGNAT addresses, not the LAN addresses,
and NTP must come up before DNS. Add a validation test asserting the
config is present and both servers actually answer DNS and NTP queries.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-28 05:23:13 -05:00
mrcharles 1044224d02 feat: wire Pi-hole to forward knel.net to local Technitium + document setup
Add a shared Docker network (dnsnet, 10.53.0.0/24) connecting Pi-hole and
Technitium containers so Pi-hole can conditionally forward knel.net and
Tailscale-reverse queries to the local authoritative Technitium instance
(10.53.0.53) instead of netboot's upstream 192.168.3.16. Also adds
end-to-end documentation for both the reference node (pfv-netboot) and
the replicated nodes (pfv-netinfra-01/02).

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-28 05:16:01 -05:00
mrcharles 291a7ebc27 feat: full re-audit of all 7 hosts with fresh ground truth
Deployed check.sh to all 7 hosts at 21:50 CDT. Captures the live state
after the user's PDM migrations:

Cnode movements since last audit:
- cnode1: tsys1 -> tsys9
- cnode2: tsys6 -> tsys7
- cnode5: tsys6 -> tsys7, storage D5(tsys4) -> S2(tsys5)

Wnode changes:
- wnode-tsys1 (102): new VM on S2, stopped
- wnode-tsys3: RAM bumped 20 -> 28 GB
- wnode-tsys6: now running (was stopped)
- wnode-tsys9: storage moved S3 -> S2

Storage distribution improved from 90/10 to 73/27 (tsys4/tsys5).
Still need 2 more cnode moves for etcd quorum survival.

Updated executive summary, k8s distribution tables, storage
utilization, and open items with the fresh data. Captured future k8s
requirements: vcluster + Rancher, OIDC to Keycloak, workload isolation
(RackRental/Suborbital ITAR/non-ITAR/SLP), and solar-aware scale-out
with PowerEdge 19xx/2950 systems.

Added tsys9 to deploy-check.sh host list.
2026-07-27 22:04:05 -05:00
mrcharles 53954d8494 docs: lock in storage philosophy and PDM migration capability
Storage philosophy (user directive):
- NVMe/SSD: k8s worker scratch + ultix-streaming (dev workstation
  running "cluster of 1" pre-prod jobs before full k8s deployment)
- Spinning rust: all infrastructure VMs (UCS, netinfra, LibreNMS, SIEM)

Clarified that hosts are standalone but managed via Proxmox Datacenter
Manager (PDM), which supports VM migration between nodes through the
UI -- eliminating the need for manual disk copies in the migration plan.

Updated all migration steps to reference PDM storage migrate instead
of manual cp commands.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 21:46:48 -05:00
mrcharles 6814374bac fix: keep UCS on spinning disk, repurpose D3 SSD for k8s/CI workloads
UCS (LDAP/AD) does not need SSD. Corrected placement:
- ucs-01 stays on D2 (tsys4 HDD) -- no change needed
- ucs-02 moves from D5 (tsys4) to S2 (tsys5 HDD) for cross-server
  redundancy only, not for speed

D3 SSD (moving to tsys5 Friday) is now freed for workloads that
actually benefit: sectestbed k8s nodes, CI/CD artifact cache, or a
container image registry mirror.

Added storage tier characterization per host (tsys4 = bulk/spinning
only, tsys5 = hybrid fast-tier hub, tsys2/3 = SSD/NVMe only). Added
UCS migration TODO item for today.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 11:51:30 -05:00
mrcharles 019ab468bf feat: relocate D3 SSD from tsys4 USB to tsys5 SAS on Friday
Deep-dive on tsys4/5 storage controllers revealed:
- tsys4: 4-port SATA, all occupied; D3 SSD is USB 2.0
- tsys5: LSI SAS1068E (8-port, 5 free) + ICH10 SATA (2 free) + USB 3.0

Decision: move D3 (SK hynix SSD) to a tsys5 SAS port. Eliminates USB
bottleneck and consolidates all fast storage on tsys5: NVMe (local) +
D3 SSD + T5-SSD + 4 HDD exports.

Updated storage architecture, speed tiers, critical VM placement,
failure survival matrix, StorageClass design, and Friday plan.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 11:50:16 -05:00
mrcharles ea2048dc87 fix: correct storage bus types after deep-dive on tsys6/7
Verified via lsblk TRAN column and dmesg that tsys6/7 local-lvm is
backed by WD My Passport portable HDDs on USB 2.0 (~30 MB/s), not
internal SATA/SAS as previously assumed. Both R610 and R620 only have
EHCI controllers -- no USB 3.0/xHCI. tsys7 has 5 empty SATA ports but
no drives installed; tsys6's SAS controller is disabled in BIOS.

Corrected all wnode recommendations: tsys6/7 stay on NFS (faster than
their USB 2.0 local-lvm). Only tsys3/5/9/2 have local storage suitable
for VM boot disks.

Also discovered tsys3 boots from Samsung PM961 NVMe 512GB (previously
documented as HDD). Its 349 GB local-lvm is the fastest existing local
storage in the fleet.

Updated speed tier table, wnode placement tables, storage distribution
matrices, migration plan, and caching strategy across PROJECT.md and
K8S.md.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 11:45:06 -05:00
mrcharles db575080c9 docs: integrate tsys2 hardware findings from Windows inventory
tsys2 validated via PowerShell on Windows: 2 SSDs (Samsung 960 PRO
NVMe 512GB + Samsung 850 EVO SATA 1TB) -- the best local storage in
the fleet. wnode-tsys2 needs no NFS; boot on NVMe, bulk data on SATA
SSD. Both NICs are USB dongles (ASIX + Realtek) -- unavoidable on
this laptop form factor, no PCIe NIC expansion possible.

Updated storage tiers, wnode placement, and k8s labeling to reflect
that tsys2 is the fastest-storage worker node. Marked TODO.md items
as resolved.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 11:39:59 -05:00
mrcharles 11ae08f7e6 docs: add TODO.md with tsys2 Windows inventory commands
Adds the PowerShell and cmd.exe commands the user needs to run on the
Windows host (pfv-tsys2) to gather disk, network, RAM, GPU, and PCIe
details before rebuilding it as Proxmox. Also documents the Friday
maintenance window steps (tsys4 NIC+RAM, tsys5 cable+NVMe) and open
questions for the next session.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 11:35:33 -05:00
mrcharles a791baa085 chore: ignore python bytecode and document git workflow
Add __pycache__/*.pyc to .gitignore and remove the tracked bytecode
file. Add version control instructions to AGENTS.md for future sessions.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 11:31:49 -05:00
mrcharles 422999bf3c chore: initialize repo with full project state
Performance optimization engagement for a 7-host Proxmox R&D cluster.
Captures the accumulated work across host tuning, network analysis,
fleet assessment, and kubernetes architecture planning.

Contents:
- Host-side tunings (scripts/): CPU governor, swappiness, BBR, NFS
  nconnect, tuned profiles -- complete on 5 of 7 hosts
- Validation + benchmarking scripts: iperf matrix, bond/NFS fixes
- Collected host data (returned-logs/): check.sh output from all 7
  hosts + iperf results, including newly-validated pfv-tsys9
- AGENTS.md: operating context for AI agents
- PROJECT.md: board-ready fleet assessment with VM placement and
  storage redundancy analysis (40 VMs across 7 hosts)
- K8S.md: kubernetes architecture deep-dive covering cnode/wnode
  distribution, StorageClass design, and ETL/HPC workload planning

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-07-27 11:31:29 -05:00
mrcharles 4201f3e669 chore: ignore framework LOGFILENAME timestamp artifacts
The framework defines LOGFILENAME as "$0.<Weekday>-YYYY-MM-DD-HH:MM:SS.$$"
and PrettyPrint appends every print_info/print_error line to it, so
executing any script that sources the framework leaves a timestamped log
file beside it (e.g. run-tests.sh.Monday-2026-07-27-10:44:31.123). These
are runtime artifacts, not source, and were showing up as untracked
noise. Ignore them across the whole repo.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 11:01:15 -05:00
mrcharles 65b972e623 fix(security): actually set Webmin 2FA directives in miniserv.conf
configure_webmin_2fa used `sed -i ... || echo ... >>` to add
twofactor_provider and twofactor to /etc/webmin/miniserv.conf. sed
returns 0 even when it matches nothing, so when the directives were
absent (the normal case on a fresh Webmin install) the `|| echo` branch
never ran. The script printed "Webmin 2FA configuration completed" while
leaving 2FA entirely unconfigured — caught by 2fa-validation reporting
"Webmin TOTP provider not configured".

Guard each directive with grep so it is appended when absent and updated
when present.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:58:10 -05:00
mrcharles bd00b61047 fix(tests): correct root/SSL/package false failures in validation suite
Three tests produced false failures when run on the deployed host:

- safe-download: the read-only-location assertion expects a write to
  fail, but the suite runs as root and root bypasses filesystem
  permissions, so the write succeeded. Skip that assertion as root.
- 2fa-validation: package presence used `dpkg -l | grep`, whose
  fixed-width output wraps long names when COLUMNS is narrow (as in a
  non-interactive shell), falsely reporting libpam-google-authenticator
  and qrencode as missing even though they were installed. Use dpkg -s.
- https-enforcement: SSL validation passed --cert-status, which requires
  OCSP stapling that many valid CDNs do not provide, flagging valid
  certificates as invalid. Drop it; --ssl-reqd still enforces TLS and
  certificate-chain verification.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:53:33 -05:00
mrcharles 1fb1413f5b fix(tests): repair test-runner arithmetic and false-positive checks
Three bugs prevented the validation suite from running cleanly:

- run-tests.sh used `((TESTS_PASSED++))` under `set -e`. Post-increment
  evaluates to the old value, so the first passing test (0 -> 1) made
  `(( ))` return 1 and errexit aborted the whole run after exactly one
  test. Use plain arithmetic assignment instead.
- https-enforcement.sh's comment filter ran `grep -n` (which prefixes
  "linenum:") and then tried to drop comment lines with
  `^[[:space:]]*#`, which never matched the line-number prefix. Every
  http:// URL in a comment (deprecated curl lines, the strict-mode
  attribution comment) was flagged as a violation. Match the prefix.
- 2fa-validation.sh hardcoded `/home/$user/` for the setup-instructions
  check, so for root it looked in /home/root (which does not exist)
  instead of /root. Resolve the home directory with getent.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:48:50 -05:00
mrcharles a54da7a43a test(validation): route post-deploy ops through guest agent for 2FA
secharden-2fa enforces AuthenticationMethods publickey,keyboard-
interactive, so once setup completes no non-interactive SSH client can
authenticate (a TOTP token is required). The harness's post-deploy steps
— log fetch, repo path resolution, and the validation suite — all relied
on SSH and therefore failed after the first successful deploy, masking
the fact that setup itself had completed (rc=0).

- remote.sh: add a vm-guest mode that runs commands as root inside the
  VM via the Proxmox qemu-guest-agent (qm guest exec), bypassing SSH/2FA
  entirely. Output is parsed on the Proxmox host with python3.
- vm-validation.sh: resolve repo path, fetch the setup log, and run the
  validation suite via vm-guest when SSH is unavailable. Detect the
  setup exit marker from the always-available live stream as a fallback
  to the fetched log. Make restore_vm_access 2FA-aware so a post-deploy
  SSH failure is understood (not a hard error) once 2FA is in effect.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:42:45 -05:00
mrcharles 40dfda47f2 fix(security): enable KbdInteractiveAuthentication for SSH 2FA
configure_ssh_2fa only enabled the deprecated ChallengeResponseAuthentication
directive (removed as a usable knob in modern OpenSSH; it no longer controls
keyboard-interactive). The base tsys-sshd-config ships
KbdInteractiveAuthentication no, so on Debian 13 (OpenSSH 9.x/10.x)
keyboard-interactive stayed disabled. With AuthenticationMethods set to
"publickey,keyboard-interactive", sshd -t then failed:

  Disabled method "keyboard-interactive" in AuthenticationMethods list ...
  AuthenticationMethods cannot be satisfied by enabled authentication methods

which aborted provisioning under errexit.

Add the modern KbdInteractiveAuthentication yes directive alongside the
legacy one so 2FA works on both current and older OpenSSH.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:37:18 -05:00
mrcharles 21cc6ee54c fix(config): newline-terminate all deployed config files
29 files under ProjectCode/ConfigFiles lacked a trailing newline. They
deploy via `cat file > target`, and several targets are subsequently
appended to (notably /etc/ssh/sshd_config, which configure_ssh_2fa
appends `AuthenticationMethods publickey,keyboard-interactive` to).
Without a trailing newline the append fused onto the last line,
producing `LoginGraceTime 60AuthenticationMethods ...`, which sshd -t
rejected as an invalid time value and aborted provisioning under errexit.

This is the same defect class that already broke the managed
authorized_keys files. Add the trailing newline to every config file
that was missing one so the cat-then-append pattern is always safe.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:35:17 -05:00
mrcharles b19bc87361 fix(security): resolve user home dir for 2FA setup instructions
setup_user_2fa wrote each user's 2FA-setup instructions to the quoted
path "~$user/2fa-setup-instructions.txt". Tilde expansion does not occur
inside double quotes, so the path was treated literally and the write
failed with "No such file or directory", aborting the whole 2FA module
(and thus provisioning) under errexit.

Resolve the home directory explicitly with `getent passwd` and use that
absolute path for both the instructions file and the chown. Skip the
user cleanly if no home directory exists.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:16:48 -05:00
mrcharles 6d77775bd6 test(validation): preserve sandbox access across SSH hardening
secharden-ssh intentionally replaces authorized_keys with the managed
production key set, which locks out the bootstrap/dev key the validation
harness uses to drive the VM. After the first deploy that reaches SSH
hardening, the harness could no longer connect to fetch logs or run the
test suite, breaking the iteration loop.

Add restore_vm_access(): after each deploy, if SSH is unreachable, it
re-injects the validation pubkey OUT OF BAND via the Proxmox guest agent
(qm guest exec runs as root inside the VM and does not depend on SSH).
The injected payload is prefixed with a newline to avoid key
concatenation when the managed file lacks a trailing newline.

Config: ACCESS_PUBKEY (default ~/.ssh/id_ed25519.pub), RESTORE_ACCESS=1.
Disable with RESTORE_ACCESS=0.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:13:04 -05:00
mrcharles 5f26f7dca1 fix(provisioning): make wazuh-agent start best-effort
secharden-wazuh.sh did a hard `systemctl start wazuh-agent`. The agent
attempts to reach its manager (tsys-nsm.knel.net) during startup; when
that host is unreachable (e.g. an isolated lab/sandbox VM, or the SIEM
being temporarily down during a fresh build), systemd's start exceeds
its timeout and the whole provisioning aborts under the framework's
errexit.

The agent is already installed and enabled, so it will keep retrying the
manager on its own. Make the start non-fatal with `|| true` so a host can
finish building even when the manager isn't reachable at deploy time.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:13:04 -05:00
mrcharles 53d953e092 fix(security): newline-terminate managed authorized_keys files
Both managed authorized_keys files lacked a trailing newline. sshd reads
the final key fine on its own, but any subsequent append (tooling, a
follow-up key, or the validation harness) concatenated onto the last key
line, fusing two keys into one unparseable blob and silently breaking
public-key auth for both.

Add the trailing newline so the files concatenate safely.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:12:55 -05:00
mrcharles 163ef9de16 fix(security): repair KexAlgorithms leading-space in sshd hardening
The KexAlgorithms line in ssh-audit-hardening.conf began with a space.
In sshd_config a leading whitespace marks a line continuation, so the
entire directive was absorbed as arguments to the (non-existent)
previous directive. The effective kexalgorithms collapsed to only the
two trailing GSSAPI entries (gss-curve25519-sha256-, gss-group16-sha512-),
which no normal OpenSSH client can negotiate.

Result: after secharden-ssh deployed this file, every SSH connection to
the host died in [preauth] with no usable key exchange algorithm. sshd -t
still returned 0, so the breakage was completely silent.

Drop the leading space so the directive is parsed as intended. This
restores normal client compatibility while keeping the hardened
algorithm set.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 10:12:38 -05:00
mrcharles a4cdd2ee30 fix(provisioning): make post-purge autoremove non-interactive
The `apt-get --purge autoremove` after the package removal pass had no
`-y`, so once removing modemmanager/wpasupplicant orphaned eleven
dependent packages, autoremove printed its "[Y/n] Abort." confirmation
and exited 1, aborting the run under errexit.

Add `-y` so the now-orphaned dependencies are purged without prompting,
consistent with the surrounding non-interactive apt invocations.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 09:58:28 -05:00
mrcharles 550d2cd078 fix(provisioning): make latencytop/cockpit-tests install best-effort
global-installPackages installed latencytop and cockpit-tests on every
non-Kali host, but both have been dropped from Debian trixie (latencytop
is dead upstream; cockpit-tests has no candidate). Under the framework's
errexit the failing apt-get aborted the entire run with rc=100 right
after the core package install completed.

These are optional monitoring/test extras, not core requirements, so
make the install best-effort with `|| true` to match the script's
existing treatment of non-critical operations.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 09:56:25 -05:00
mrcharles 83e4d7e8ce test(validation): add reproducible git-based VM validation harness
Add an end-to-end validation loop for the sandbox VM and the single
SSH/SCP chokepoint it depends on:

- Project-Tests/remote.sh: the only place ssh/scp is invoked. Provides
  prox / vm / vmroot / *-file / *-copy modes. Centralizes host/user/key
  config and keeps remote access auditable and reusable.
- Project-Tests/vm-validation.sh: drives a Proxmox VM through
  snapshot -> deploy -> validate with one-command rollback. Deployment is
  GIT-BASED: the VM clones/pulls the public repo itself, exactly as a
  fresh server would, so results are identical regardless of who runs it
  (no reliance on a local working copy or rsync). Resolves the absolute
  repo path on the VM before sudo to avoid the '~' -> root's home trap.
- logs/.gitignore: ignore generated validation/test logs (was a no-op
  `!.gitignore` with no ignore rule; logs would have been committed).

Also fixes a `help`-branch typo (`${BASH_SOURCE[0]}`) and adds
PROX_USER (defaults to root) since the bare Proxmox hostname defaulted
to the wrong SSH user.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 09:32:14 -05:00
mrcharles 377c83bcf1 fix(framework): guard tput calls against missing TERM
The framework sources ErrorHandling.sh which enables `set -o errexit`
globally, yet PrettyPrint's print_info/print_error invoked `tput bold`
and `tput sgr0` with no protection. In any TERM-less context (SSH
automation, CI, cron) `tput` fails with "unknown terminal" and, under
errexit, aborts the entire script on the very first status message.

Suppress tput stderr and add `|| true` so the color helpers degrade
gracefully to plain output instead of crashing every consumer script.
This is an internal framework consistency fix: strict mode + unguarded
external command were mutually incompatible.

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-27 09:32:04 -05:00
Charles N Wyble 97ff9c321d docs(agents): document repo layout and autonomous git policy
Make AGENTS.md actionable for future agents by recording what was learned
while fixing the provisioning scripts:

- Add a Repository Layout section: the KNELShellFramework is vendored
  under vendor/.../KNELShellFramework (not at repo root), scripts must
  self-locate via BASH_SOURCE, configs are read locally (no CDN), and
  some .sh agents are actually PHP
- Replace the vague "commit immediately" note with an explicit
  Autonomous Git Workflow section authorizing agents to commit AND push
  without being asked, grouped into coherent atomic commits

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-25 13:49:55 -05:00
Charles N Wyble 5928d96aec fix(tests): repair framework paths, assertions, and arithmetic
The test suite referenced a Framework-Includes/ directory that does not
exist at the repo root (it is vendored under vendor/.../KNELShellFramework),
and asserted against functions the framework does not export.

- Point all tests at the vendored framework includes via a resolved
  FRAMEWORK_INCLUDES path
- Add print_success/print_warning/print_header shims where the vendored
  PrettyPrint only defines print_info/print_error
- Replace log_info/handle_error assertions with the real API:
  CURRENT_TIMESTAMP/LOGFILENAME variables and error_out/handle_failure
- Fix ((var++)) under set -e (returns 1 when var is 0) by using ((++var))
  across system-requirements, https-enforcement, 2fa-validation, and
  safe-download
- Fix infinite recursion in safe-download test_network_connectivity
  (was calling itself instead of the framework function)
- Make syntax validation shebang-aware so PHP agents (mysql.sh) are
  skipped instead of flagged as bash syntax errors

Framework unit test now passes (exit 0).

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-25 13:49:47 -05:00
Charles N Wyble 688b7190e6 refactor(provisioning): make scripts self-locating and read configs locally
All provisioning scripts and modules now resolve their own location via
BASH_SOURCE and derive PROJECT_ROOT_PATH from it, removing a hard
dependency on the current working directory.

- Derive PROJECT_ROOT_PATH/CONFIGFILES_PATH/MODULES_PATH/SCRIPTS_PATH
  from BASH_SOURCE in SetupNewSystem.sh and every module
- Replace all curl ${DL_ROOT}/... downloads with cat of the matching
  local files under ProjectCode/ConfigFiles (the dl.knownelement.com CDN
  is no longer required for a git clone)
- Invoke modules by absolute path instead of cd ./Modules/X && bash ./x
- Fix secharden-audit-agents.sh: wrong path depth (../../ vs ../../..),
  wrong Project-Includes glob, and ConfigFiles/AudidD -> AuditD typo,
  all of which previously crashed the script
- Remove duplicate FrameworkVars source lines

Run from anywhere with: sudo bash ProjectCode/SetupNewSystem.sh

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-07-25 13:49:39 -05:00
Charles N Wyble 65d4985d16 docs: add AGENTS.md with git commit guidelines
Add agent guidelines for AI assistants working on this repository:

- Document atomic commit requirements
- Specify conventional commit format with examples
- Require verbose, formatted commit messages
- Emphasize immediate commit/push behavior

🤖 Generated with [Crush](https://github.com/charmassociates/crush)

Assisted-by: GLM-5 via Crush <crush@charm.land>
2026-02-17 17:08:44 -05:00
mrcharles 453f3a22ac a few minor annoyances now addressed 2025-12-29 14:17:08 -05:00
mrcharles f283f8cfb8 a bit of refactoring 2025-12-28 18:49:30 -05:00
mrcharles d33c8df277 AI project review 2025-12-28 18:46:53 -05:00
mrcharles 2930eeaf27 ssh pub key regression, need to use cat instead of curl 2025-07-29 13:54:39 -05:00
mrcharles 870540840c use the tailscale installer 2025-07-29 13:45:21 -05:00
mrcharles 5e2eaff55d typo 2025-07-29 13:42:14 -05:00
mrcharles 8f19c9fb6e kali corner case... 2025-07-29 13:32:53 -05:00
mrcharles 40ab4608e2 some minor ubuntu default cleanup 2025-07-17 23:04:40 -05:00
mrcharles 47ddb93fef Adding ansible-core to be able to run compliance as code playbooks 2025-07-16 09:37:11 -05:00
mrcharles e73b81e229 . 2025-07-14 13:08:05 -05:00
mrcharles 39e37d0f76 . 2025-07-14 13:04:31 -05:00
mrcharles 31e66864ad . 2025-07-14 13:02:42 -05:00
mrcharles 0006eefcf1 . 2025-07-14 12:58:25 -05:00
mrcharles abfaf765e6 . 2025-07-14 12:55:48 -05:00
mrcharles 1f2bd31380 . 2025-07-14 12:53:41 -05:00
mrcharles 93cea874a8 . 2025-07-14 12:50:48 -05:00
mrcharles a898ebc59d . 2025-07-14 12:49:26 -05:00
mrcharles 78cc8cbcf3 . 2025-07-14 12:47:40 -05:00
mrcharles 495d0bb03b . 2025-07-14 12:46:53 -05:00
mrcharles 7a7d23f36c . 2025-07-14 12:42:22 -05:00
mrcharles 84f3ca3b0e . 2025-07-14 12:38:07 -05:00
mrcharles f9f32612bb . 2025-07-14 12:37:04 -05:00
mrcharles 09063bfee4 case matters... 2025-07-14 12:36:03 -05:00
mrcharles 5bbaff89e9 refactored to use vendored shell framework. lets test. 2025-07-14 12:34:33 -05:00
mrcharles 5a8561ea84 Update "KnelShell" from "ssh://git@git.knownelement.com:29418/KNEL/KNELShellFramework.git@main"
git-vendor-name: KnelShell
git-vendor-dir: vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework
git-vendor-repository: ssh://git@git.knownelement.com:29418/KNEL/KNELShellFramework.git
git-vendor-ref: main
2025-07-14 12:18:27 -05:00
mrcharles 2fa32a5eb7 Squashed 'vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework/' changes from 5ecde81..1fb5a06
1fb5a06 Added SafeDownload and added shebang to DebugMe

git-subtree-dir: vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework
git-subtree-split: 1fb5a061aecd61df406e5ebcdb010097f1ccbc69
2025-07-14 12:18:27 -05:00
mrcharles 83d5cf2f8d moved docs
Switching to using vendored shell framework
moved SafeDownload to vendored shell framework repo
2025-07-14 12:17:29 -05:00
mrcharles 49e57ff846 Squashed 'vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework/' content from commit 5ecde81
git-subtree-dir: vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework
git-subtree-split: 5ecde81ce441d5802fd7e7e91a441e34f327f457
2025-07-14 12:11:42 -05:00
mrcharles 47b5a976c2 Add "KnelShell" from "ssh://git@git.knownelement.com:29418/KNEL/KNELShellFramework.git@main"
git-vendor-name: KnelShell
git-vendor-dir: vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework
git-vendor-repository: ssh://git@git.knownelement.com:29418/KNEL/KNELShellFramework.git
git-vendor-ref: main
2025-07-14 12:11:42 -05:00
mrcharles a710fc7b4e removed debugging bits 2025-07-14 11:04:21 -05:00
mrcharles c6e458de8b . 2025-07-14 11:03:08 -05:00
mrcharles e31bab4162 . 2025-07-14 11:01:19 -05:00
mrcharles 86740b8c7d . 2025-07-14 10:59:32 -05:00
mrcharles f585f90b7f . 2025-07-14 10:55:54 -05:00
mrcharles 24c10b6f35 it hallucinated print_header 2025-07-14 10:50:42 -05:00
mrcharles 634a998d7e testing 2025-07-14 10:48:59 -05:00
mrcharles e3685f68ad forgot to call the function 2025-07-14 10:33:04 -05:00
mrcharles ac857c91c3 actually run the 2fa script. 2025-07-14 10:31:22 -05:00
mrcharlesandClaude a632e7d514 Implement comprehensive two-factor authentication for SSH and web services
- Complete rewrite of secharden-2fa.sh with full 2FA implementation
- SSH 2FA using Google Authenticator with publickey + TOTP authentication
- Cockpit web interface 2FA with custom PAM configuration
- Webmin 2FA support with automatic detection and configuration
- User setup automation with QR codes and backup codes generation
- Gradual rollout support using nullok for phased deployment
- Automatic configuration backup and restore procedures
- Add 2fa-validation.sh security test for comprehensive validation
- Create TSYS-2FA-GUIDE.md with complete implementation documentation
- Add DEVELOPMENT-GUIDELINES.md with coding standards and best practices
- Optimize package installation with single apt-get commands for performance

The 2FA implementation provides enterprise-grade security while maintaining
usability and proper emergency access procedures. Includes comprehensive
testing, documentation, and follows established security best practices.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-14 10:23:07 -05:00
mrcharlesandClaude f6acf660f6 Implement comprehensive testing framework and enhance documentation
- Add Project-Tests directory with complete testing infrastructure
- Create main test runner with JSON reporting and categorized tests
- Implement system validation tests (RAM, disk, network, permissions)
- Add security testing for HTTPS enforcement and deployment methods
- Create unit tests for framework functions and syntax validation
- Add ConfigValidation.sh framework for pre-flight system checks
- Enhance documentation with SECURITY.md and DEPLOYMENT.md guides
- Provide comprehensive testing README with usage instructions

The testing framework validates system compatibility, security configurations,
and deployment requirements before execution, preventing deployment failures
and providing clear error reporting for troubleshooting.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-14 09:35:27 -05:00
mrcharlesandClaude 0c736c7295 Enforce HTTPS for all downloads to eliminate security vulnerabilities
- Convert 16 HTTP URLs to HTTPS across 3 critical scripts
- Dell OMSA script: Ubuntu archive and Dell repository URLs now use HTTPS
- Proxmox legacy script: Download URLs converted to secure connections
- SSL stack script: Apache source URLs updated to official archive
- Update documentation to reflect resolved security issues
- Mark HTTPS enforcement as completed in todo lists

This addresses the second critical security concern from the security review,
eliminating man-in-the-middle attack vectors during package downloads.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-14 09:22:32 -05:00
mrcharles 273e7fe674 Claude code review of my work. 2025-07-12 00:17:52 -05:00
mrcharles 6609d7d9e3 sigh. 2025-07-11 11:52:28 -05:00
mrcharles 0588b2dd60 ifdev for dev boxes, they have less hardened ssh config because vscode remote etc 2025-07-11 11:48:53 -05:00
mrcharles f399308b2d allow root to login to cockpit 2025-07-10 10:47:21 -05:00
mrcharles 45b53efe11 working on v1.1, secrets management/bootstrap 2025-07-10 10:28:00 -05:00
360 changed files with 128761 additions and 538 deletions
+3
View File
@@ -0,0 +1,3 @@
*
!crush.json
!.gitignore
+11
View File
@@ -0,0 +1,11 @@
{
"hooks": {
"PreToolUse": [
{
"matcher": "^(bash|edit|write|multiedit|lsp_replace_symbol|lsp_rename)$",
"command": "./hooks/ticket-gate.sh",
"timeout": 5
}
]
}
}
+7
View File
@@ -0,0 +1,7 @@
# PFVCluster environment variables
# Copy to .env and fill in values for local development/testing.
# Pi-hole (netinfra/pihole/docker-compose.yml)
PIHOLE_WEB_PASSWORD=changeme
# Shellcheck wrapper (tests/shellcheck.sh) — no config needed, uses Docker.
+46
View File
@@ -0,0 +1,46 @@
# Crush internal state (track crush.json config, ignore session data)
.crush/crush.db
.crush/crush.db-*
.crush/logs/
.crush/memory/
.crush/active-ticket
# OS/editor
.DS_Store
*.swp
*.swo
*~
\#*\#
.\#*
# Python bytecode
__pycache__/
*.pyc
# Temporary files
*.tmp
*.bak
*.log.tmp
# LOGFILENAME artifacts: the framework (Logging.sh + PrettyPrint.sh) appends
# every print_info/print_error line to LOGFILENAME, defined as
# "$0.<Weekday>-YYYY-MM-DD-HH:MM:SS.$$". Running any script that sources the
# framework therefore drops a timestamped log file next to it.
*.Monday-*
*.Tuesday-*
*.Wednesday-*
*.Thursday-*
*.Friday-*
*.Saturday-*
*.Sunday-*
# Sensitive exports / runtime data
dns-cluster-setup/.export/
returned-logs/
# Kubernetes secrets (kubeconfig contains embedded client certs)
k8s/kubeconfig.yaml
k8s/*.token
# Pi-hole web UI password (real value only in on-box .env, never committed)
netinfra/pihole/.env
+17
View File
@@ -0,0 +1,17 @@
# ShellCheck configuration for PFVCluster
# (used when running `shellcheck` directly; tests/shellcheck.sh applies the
# same disables via -e for consistent results under Docker)
#
# These checks are DISABLED because they flag intentional conventions of this
# codebase, not bugs:
#
# SC1090 / SC1091 — cannot follow dynamically-computed `source` paths. The KNEL
# framework (vendor/) and test harness source helpers via computed include
# dirs, which shellcheck cannot resolve statically.
# SC2029 — ssh orchestration (tests/remote.sh and perf/k8s/dns scripts)
# deliberately builds and expands the remote command on the CLIENT side before
# sending it. That is the whole point of the single-chokepoint remote pattern.
disable=SC1090,SC1091,SC2029
# Treat external-sourced files as bash (matches #!/usr/bin/env bash framework).
external-sources=true
-6
View File
@@ -1,6 +0,0 @@
{
"debug.javascript.defaultRuntimeExecutable": {
"pwa-node": "/home/localuser/.local/share/mise/shims/node"
},
"python.defaultInterpreterPath": "/home/localuser/.local/share/mise/installs/python/3.11.13/bin/python"
}
+139
View File
@@ -0,0 +1,139 @@
# ADOPTING.md — How to adopt this framework into an existing project
> This is the guide you give to an agent (or human) that says:
> "Look at `~/daytoday/meta` and adopt its rules/patterns for this project."
>
> The framework lives at: `ssh://git@git.knownelement.com:29418/TSYSGroupCorporate/TSYSGroupAIOS.git`
> Template repo: https://git.knownelement.com/TSYSGroupCorporate/TSYSGroupAIOS
---
## For a NEW project (greenfield)
```bash
# Create from the Gitea template, clone, done:
tea repo create --owner <org> --name <project> --template-from TSYSGroupCorporate/TSYSGroupAIOS
# or clone directly:
git clone ssh://git@git.knownelement.com:29418/TSYSGroupCorporate/TSYSGroupAIOS.git <project>
cd <project>
make setup # install git hooks (= bash scripts/setup-hooks.sh)
make fast # verify baseline (= bash scripts/check-rules.sh --fast)
```
Then edit `AGENTS.md` (fill bracketed fields), override `scripts/test.sh`, and start work.
---
## For an EXISTING project (brownfield — e.g. PFVCluster)
Adoption is incremental. You don't rewrite the project — you overlay the
framework's enforcement layer and adjust the project's existing conventions to
match. Do these steps in order:
### Step 1: Copy the enforcement layer
```bash
cd ~/projects/<existing-project>
# Bring in the framework files that don't already exist.
# Get them from the meta repo:
META=~/daytoday/meta
# Scripts, hooks engine, shared lib (the mechanical enforcement)
cp -n "$META/Makefile" .
mkdir -p scripts/lib
cp -n "$META/scripts/check-rules.sh" scripts/
cp -n "$META/scripts/setup-hooks.sh" scripts/
cp -n "$META/scripts/pre-commit" scripts/
cp -n "$META/scripts/pre-push" scripts/
cp -n "$META/scripts/docker-run.sh" scripts/
cp -n "$META/scripts/garden.sh" scripts/
cp -n "$META/scripts/lib/common.sh" scripts/lib/
chmod +x scripts/*.sh scripts/pre-commit scripts/pre-push
```
### Step 2: Bring in the workflow files (if the project doesn't have them)
```bash
cp -n "$META/STATUS.md" . # if the project's STATUS.md is a Discourse pointer, KEEP IT
cp -n "$META/WORKING.md" .
cp -n "$META/questions-v1.md" .
cp -n "$META/.env.example" . # only if the project doesn't have one
```
**Important:** if the project already uses Discourse as its SoR (like PFVCluster),
its existing `STATUS.md` may be a pointer stub to Discourse. In that case, do NOT
overwrite it — the project already has the right pattern. The template `STATUS.md`
is for projects that don't have one yet.
### Step 3: Install git hooks and verify
```bash
bash scripts/setup-hooks.sh # installs pre-commit + pre-push from scripts/
bash scripts/check-rules.sh --fast # see what passes and what fails
```
### Step 4: Fix the failures (incrementally)
`bash scripts/check-rules.sh --fast` will likely report failures — the project's existing code may not
pass shellcheck, or `.md` files lack Discourse pointers. **Fix these
incrementally; don't rewrite the project in one pass.**
Common fixes:
- **shellcheck violations:** run `bash scripts/check-rules.sh` for details; fix warnings in the flagged files.
- **`:latest` image tags:** pin to a specific version in docker-compose / Dockerfile.
- **Container naming:** add `container_name:` to every service in docker-compose files.
- **Missing required files:** create `questions-v1.md`, `.env.example`, etc.
- **Discourse pointer-header:** for `.md` files that should be Discourse stubs, migrate content to Discourse and leave a pointer. For operational files (`AGENTS.md`, `STATUS.md`, etc.), add them to `PROJECT_DOC_EXEMPT`.
### Step 5: Merge the project's AGENTS.md with the template
Read both the project's existing `AGENTS.md` and the template's (`~/daytoday/meta/AGENTS.md`).
Merge by:
1. Keeping all project-specific content (VM paths, auth details, domain knowledge).
2. Adding the template's standard sections the project is missing (Quick Start, Systems of Record, Working Style, Key Commands, Enforcement Model).
3. Replacing any conflicting policy with the baseline (the template wins on cross-project conventions; the project wins on domain specifics).
### Step 6: Configure project-specific env vars
Set these in the project's `.env` or in the Makefile to customize checks:
```bash
PROJECT_DOC_EXEMPT="AGENTS.md STATUS.md WORKING.md ..." # files that don't need Discourse pointers
PROJECT_DISCOURSE_HOST="community.turnsys.com" # Discourse instance
PROJECT_REQUIRED_FILES="..." # extra required files beyond the defaults
PROJECT_BANNED_SUFFIXES="py|js|ts" # banned production file types (optional)
```
### Step 7: Commit and push
```bash
bash scripts/check-rules.sh # full audit should pass
git add -A
git commit -m "chore: adopt TSYSGroupAIOS framework (git hooks, rules engine, SoR policy)"
git push
```
---
## What NOT to change during adoption
- **Don't rewrite existing code** that works. The framework enforces conventions going forward; fix existing violations incrementally via `bash scripts/check-rules.sh`.
- **Don't remove the project's Redmine/Discourse integration.** The framework *requires* it — the project already has it. Align the AGENTS.md prose to match.
- **Don't add `docs/JOURNAL.md`.** Redmine is the system of record for work; Discourse for docs. No JOURNAL.md.
- **Don't add Crush hooks.** The framework is harness-agnostic. Enforcement is git hooks + AGENTS.md prose only.
---
## Quick reference: what the framework gives you
| What | Files | Portable? |
|---|---|---|
| Git hooks (pre-commit/pre-push) | `scripts/pre-commit`, `scripts/pre-push` | Yes — any git, any agent |
| Rules engine | `scripts/check-rules.sh` | Yes |
| Shared bash library | `scripts/lib/common.sh` | Yes |
| Docker wrapper | `scripts/docker-run.sh` | Yes |
| Lifecycle scripts | `scripts/up.sh`, `scripts/down.sh` | Yes |
| Gardening | `scripts/garden.sh` | Yes |
| Policy document | `AGENTS.md` | Yes — any agent framework reads it |
| Global baseline | `BASELINE-PROMPT.md` | Yes — paste into any system prompt |
+859
View File
@@ -0,0 +1,859 @@
# Agent Guidelines
> CREDS-PATTERN-RETIRED 2026-09-06 (founder): all secrets live in the TSGCOO Bitwarden vault via the KNELSecretsManager Go CLI — `sm env creds/<name>` (see /data2/TSGCOO/.local/bin/sm). Command examples below showing `--env-file ~/.creds/...` are HISTORICAL; translate to `sm env` before reuse. Cloudron account identities (ic-builder/ic-reviewer/manager-tsg/vptechops) author all work — never the founder account.
## Agent Authority (NON-NEGOTIABLE)
**No work is permissible on any system without an approved Redmine ticket.
There are no exceptions to this rule.**
### Scope of authority
1. **Tickets govern all work.** The agent performs ONLY the work described in
the approved Redmine ticket. Anything outside that scope — no matter how
small, helpful, or "obvious" — is prohibited.
2. **No autonomous system changes.** The agent does not modify, configure,
create, delete, or grant anything on a production system unless it is
explicitly directed by an approved ticket. "Production system" means
every system in the fleet — there is no "test" exception unless the ticket
says so.
3. **Propose, never implement.** If the agent discovers additional work that
should be done — a bug, a misconfiguration, a missing dependency, an
enhancement — it does NOT implement it. Instead, it creates a Redmine
sub-ticket (status Feedback) describing the finding and surfaces it to
the user for approval.
4. **Security and access changes require extra scrutiny.** Changes to sudoers,
SSH keys, user accounts, firewall rules, authentication policy, file
permissions, or any privilege-related configuration are treated as
policy decisions, not implementation details. The agent may suggest
them but NEVER implements them without explicit user direction in the
ticket or a sub-ticket the user has approved.
5. **The user makes policy. The agent implements policy.** The agent does
not decide who gets sudo, what keys go where, what services run, or what
the access model is. The agent executes the user's decisions, exactly
as specified.
6. **When in doubt, ask.** If the ticket is ambiguous, if a task seems to
require something not explicitly authorized, or if the agent is unsure
whether an action is in scope — STOP and ask the user via the ticket
or directly. Asking is always acceptable. Overstepping is never
acceptable.
7. **NEVER close a ticket without explicit user permission.** You may
SUGGEST a close when the result is clearly scoped and delivered. If
it's ambiguous whether the work is truly complete, don't suggest a
close — leave that decision to the user. This applies to ALL tickets,
no exceptions.
8. **User acceptance testing is MANDATORY before declaring work done.**
The agent performs implementation and technical validation (services
running, configs correct, APIs responding). The user performs UAT —
visually confirming dashboards render data, alerts deliver, tools are
usable. The agent MUST NOT set done-ratio to 100%, MUST NOT suggest
closing, and MUST NOT move to the next ticket until the user explicitly
accepts the work. "Technically wired but blank dashboard" is NOT done.
9. **NEVER access a database directly if an API exists.** APIs are the
stable contract; databases change schemas without warning. If a tool
has an API, use it — exhaust all API endpoints, check the docs, try
alternative methods. Only fall back to direct DB access as a last
resort AND with explicit user approval for that specific instance.
Soon all DB access will route through a proxy under zero trust;
building API-first habits now ensures that transition is clean.
### Access-channel policy: SSH only (NON-NEGOTIABLE)
**The qemu guest-agent is NEVER an access, execution, or key-delivery
channel.** SSH (`sshd`) is the only approved remote access path. Every
command must flow through sshd so it is captured by the standard
auth/logging/audit infrastructure. This is an ITAR/CMMC/TS/SCI
environment — there is no back-door exception, ever.
- **Forbidden:** `qm guest exec` (runs arbitrary commands inside a guest
over an unaudited channel) and any wrapper around it (e.g. a `vm-guest`
mode). This includes using guest-agent to *deliver* an SSH key, even if
the resulting SSH login is itself audited — the delivery bypassed audit.
- **Allowed:** installing or checking qemu-guest-agent for its intended
purpose — letting Proxmox see guest state (`qm guest cmd <id> ping`,
`agent: 1` config, `apt install qemu-guest-agent`). Visibility only;
never execution.
- **Enforced mechanically:** `scripts/check-rules.sh` rule #11 fails on
any `qm guest exec` / `vm-guest` pattern in code. `tests/remote.sh` has
no guest-exec mode.
- **If a system is locked out** (no SSH key, no guest-exec path): surface
it to the user. Do NOT improvise an alternate back-channel. The user
authorizes the unblock method (console login, credential, etc.).
### Ticket-closing policy (NON-NEGOTIABLE)
**NEVER close a Redmine ticket without explicit user permission.** You
may SUGGEST a close when the result is clearly scoped and delivered. If
it's ambiguous whether the work is truly complete, don't suggest a close
— leave that decision to the user. This applies to ALL tickets, no
exceptions.
### What this means in practice
- Discovered a typo in a config during approved work? **Finish the approved
work. Create a sub-ticket for the typo. Do not fix it inline.**
- Think a system should also have localuser sudo configured? **Do not add
it. Propose it in a sub-ticket.**
- Need to install a package the ticket didn't mention? **Ask first.**
- Found a security issue? **Create a ticket immediately with full details.
Do not remediate without approval.**
This environment operates in ITAR/CMMC/TS/SCI space. Every action must be
traceable to an approved ticket. There is no "I thought it would help."
## Quick Start
**You are an AI agent working on this project. Your first actions, in order:**
> **SESSION-START GATE (NON-NEGOTIABLE):** Steps 1-3 orient you. Step 4 is the
> check-for-understanding gate (rule summary). Steps 5-7 gather state. Step 8 is
> the scope-alignment gate. You MUST NOT begin any task work until the user
> (a) confirms your rule summary AND (b) names the ticket to work on. This runs
> **every session, automatically** — the user should never have to ask for it.
1. **Set up the environment:** `bash scripts/setup-hooks.sh` (installs git hooks — idempotent).
2. **Read this file** (`AGENTS.md`) — project policy and domain knowledge.
3. **Read the latest questions file** (`questions-v*.md`) — open questions awaiting human input. The version number increments each round (v1, v2, v3...).
4. **Check for understanding — GATE.** In your own words, summarize ALL the rules
back to the user before doing any work: Agent Authority (ticket-governed,
propose-never-implement), Access-Channel SSH-only policy, Remote access
(remote.sh chokepoint mandatory, DNS names only — never IP literals),
Questions policy (no harness question-tools), Documentation policy
(Discourse is SoR), Redmine tracking, Git policy (always commit+push,
shellcheck), Mandatory infra-change documentation (#298 audit log), SSH
routing chokepoints (`tests/remote.sh`), the mechanically enforced
`check-rules.sh` rules, and Credentials (Vault migration TODO). Then stop
and wait for the user to confirm. This checkpoint guarantees every
session starts aligned.
5. **Check Redmine**`docker run --rm --env-file ~/.creds/redmine.env git.knownelement.com/reachableceo/redmine-cli:latest list --assigned-to-me -p 55` for active work.
6. **Check current state:** `git log --oneline -10`.
7. **Run rule audit:** `bash scripts/check-rules.sh --fast`.
8. **Scope-alignment — GATE.** Present the session handoff's "What's Left"
priority list (or the Redmine queue if no handoff exists). Do NOT scan the
full ticket queue and pick work on your own — **the user directs what gets
worked on, always.** Note any new or urgent items from the Redmine check,
then ask which ticket to work on tonight. **STOP and wait.** Do not set
`.crush/active-ticket`, do not read systems, do not run diagnostics, until
the user names the target. This gate prevents the agent from burning
context on work the user didn't ask for.
## Enforcement Model
Git hooks (`scripts/pre-commit`, `scripts/pre-push`) enforce the rules defined in
`scripts/check-rules.sh`. The rules engine checks: shellcheck (zero warnings
including info-level), Docker image pinning (no `:latest`), container naming,
required files, Discourse pointer headers, and more. Run `bash scripts/check-rules.sh`
for a full audit or `--fast` for pre-commit speed. Bypass with `--no-verify`
(emergencies only).
## Task Tracking
- **Redmine is the system of record for all work.**
- **NEVER close a ticket without explicit user permission.** Suggest a
close when clearly scoped/delivered; if ambiguous, don't suggest.
- **Ticket-first enforcement (mechanically enforced).** Before starting
any work, set the active ticket: `echo '#NNN' > .crush/active-ticket`.
The Crush hook (`hooks/ticket-gate.sh`) blocks modifying operations
until this file exists. If no ticket exists, CREATE ONE FIRST via
redmine-cli, then set it. Clear when done: `> .crush/active-ticket`.
**Parallel sessions:** each session uses its OWN suffixed file (e.g.
`.crush/active-ticket-plant`, `.crush/active-ticket-core`) so concurrent
sessions never clobber each other's ticket; any non-empty
`.crush/active-ticket*` satisfies the gate.
- **WORKING.md** is the only in-repo task tracker — a scratchpad for the current
session. The pre-commit hook blocks commits while any task remains unchecked.
- **Parallel sessions:** when two sessions share this repo (e.g. physical plant
vs core infra), the second session uses a gitignored scratch tracker such as
`.crush/WORKING-PLANT.md` instead of WORKING.md, so the two trackers (and
their pre-commit gates) never collide. Wipe at session close.
- Clear WORKING.md before responding to the user.
## Rolling HUD (session-scoped, NOT persisted)
The rolling HUD is a **live status display** the agent maintains throughout
the session to help the human follow along. It is NOT a system of record —
Redmine, Discourse, and git are the durable systems. The HUD exists purely
for the human's situational awareness during the session.
- **Format:** a compact block shown at the end of each significant response
(after completing a step, hitting a blocker, or pivoting). Example:
```
┌─ SESSION HUD ────────────────────────────────────────
│ Active: #343 (Monitoring coverage matrix)
│ Done: ✓ scope-alignment gate added to both AGENTS.md
│ ✓ committed + pushed (6a2550b)
│ Now: drafting coverage matrix on Discourse
│ Next: → #341 TEMPer USB (deploy on pfv-tsys1)
│ → #338 LibreNMS alerts (BLOCKED: pushover.env)
├─ USER ACTION ITEMS ──────────────────────────────────
│ • Populate ~/.creds/pushover.env (unblocks #338, #428)
│ • Populate ~/.creds/prometheus.env + grafana.env (#430)
└──────────────────────────────────────────────────────
```
- **Placement:** may be written to `.crush/hud.md` on disk to keep context
window smaller (re-read and update rather than hold in memory). Never
committed to git. Wiped at session end.
- **Sections:**
- **Active:** current ticket number + one-line description
- **Done:** ✓ items completed this session (append as work progresses)
- **Now:** what the agent is actively doing
- **Next:** the queued items (per handoff priority or user direction)
- **User action items:** things ONLY the user can do (populate creds,
physical work, manual deploys) with the tickets they unblock
- **When to show it:** after each logical unit of work, at blockers, and
when pivoting between tickets. Not every trivial response — use judgment.
- **NOT a substitute for Redmine/Discourse/git.** The HUD is ephemeral. When
work completes, update the durable systems (ticket notes, Discourse wiki,
commits). The HUD just tracks the live narrative for the human.
## Working Style
- **Stop over-thinking.** Get to code and output faster. Explore with code;
gather ground truth. Do not burn tokens reasoning about things a quick command
answers.
- **Prefer Unix utilities** (awk, sed, grep, cut, tr) for file editing and
text processing over harness edit tools when feasible — deterministic and
exact where harness editors can be whitespace-fragile.
- **Farm work out to deterministic tooling:** linters, LSPs, formatters, test
runners. If an LSP is wired up, use it; otherwise pull a Docker image and lint
inside it.
- **Use sub-agents as subcontractors:** scoped spec in, distilled deliverable out.
Never read 10+ files sequentially; batch into agent calls.
- **Command timeouts (NON-NEGOTIABLE):** Every command that touches a remote
system MUST be wrapped with `timeout`. Hard limits: 30s for quick reads
(status, ps, ls), 120s for standard operations, 300s for deployments/pulls.
If a command hits the timeout, STOP and investigate root cause — never
blindly retry. A hung command is a failed command. Detect failure fast,
diagnose, fix, move on. Example: `timeout 120 bash tests/remote.sh vm 'cmd'`.
This applies to ALL tools — bash, docker, CLIs, sub-agents.
## Field Lessons (append at session close)
Hard-won operational gotchas. One line each; keep them short and load-bearing.
- **CLI body scanner workaround:** inline `-b "body"` arguments containing
words like `systemctl` are blocked by the command scanner. Write long bodies
to `/tmp/note.md` and pass `-b "$(cat /tmp/note.md)"` (redmine-cli +
discourse-cli both).
- **python3-venv trap:** `python3 -c 'import venv'` succeeds even without the
package; the resulting venv has no pip. Guard with `dpkg -s python3-venv`,
verify `test -x venv/bin/pip`, delete pip-less venvs and recreate.
- **temperusb API:** there is no `Temper` class. Use
`TemperHandler().get_devices()[i].get_temperatures(sensors=[0,1])`
(sensor 0 = internal, 1 = external probe).
- **LibreNMS API:** device add requires `"version": "v2c"` (with the v);
`discovery.php`/`poller.php` must run as the `librenms` user; SNMP
community fields are redacted in API responses by design.
- **remote.sh modes:** `vm-copy`/`prox-copy` destination must be a FULL file
path, never a directory; `vm-file`/`prox-file` push and execute a script
via `bash -s` (use for idempotent deploy scripts).
- **Remote restart verification:** after restarting services (especially in
parallel across hosts), re-verify state several seconds later. Immediate
checks race the restart and produce false failures.
- **NUT:** `upsd.conf` LISTEN edits only take effect after
`systemctl restart nut-server`. udev rule changes need
`udevadm control --reload-rules && udevadm trigger`.
- **Tailscale app-connector sources (HA):** HA polls originate from its LAN
IP (192.168.3.12), not its Tailscale IP. Include the HA LAN IP in every
allowlist (snmpd `rocommunity` source lists, upsd LISTEN binds).
- **`git mv` fails on untracked files:** plain `mv` first, `git add` after.
- **SNMP extend with UTF-8:** net-snmp CLI prints `°C` extend output as
Hex-STRING, but pysnmp/HA receive the decoded string. Parse the decoded
lm-sensors text lines with a regex value_template.
- **pfv-bms (Home Assistant):** no SSH by design; everything goes through the
REST API (`/api/`, `/api/config`, `/api/services`) with the long-lived
token in `~/.creds/homeassistant.env`. Instance unit system is °F — set
explicit `unit_of_measurement: "°C"` on custom sensors.
- **HAOS core_ssh on 22222 is the sanctioned host shell:** `VM_IP=pfv-bms-lan.knel.net
VM_USER=root VM_PORT=22222 bash tests/remote.sh vm '...'` — use the LAN name
(knel.net zone is Tailscale-IP based; TS inbound to pfv-bms is unreliable).
Host OS has NO git/ssh/curl — run git against /mnt/data/supervisor/homeassistant
via `docker run homeassistant/amd64-addon-git_pull:<ver>`.
- **Git pull add-on (core_git_pull) wipes /config on first run** if `.git` is
absent (rm -rf + fresh clone). ALWAYS pre-seed /config as a git checkout
(deploy key ~/.creds/pfv-bms-gitpull, repo KNEL/pfv-bms) before starting it.
- **HA core boot death-spiral:** supervisor watchdog SIGTERMs slow boots
(docker logs: "finish process received signal 15" while supervisor logs
"Still waiting for Core to start"). Fix from the 22222 shell:
`ha ha options --watchdog=false`, start core, diagnose, re-enable once stable.
- **Technitium on netinfra-02 (tsys-dns container) does not reload rsync-synced
zone files** — `docker restart tsys-dns` after record adds until #469 (native
clustering) lands. netinfra-01 (the dns-cli target) is fine.
- **Supervisor add-on options via API:** docker cp a script into
hassio_supervisor, read token from /data/cli.json `access_token`, POST
/addons/<slug>/options (the /apps/ spelling 404s).
- **NEVER restart/kill BOTH members of a redundant pair simultaneously**
(founder ruling 2026-09-01 after the DNS outage): restart/redeploy ONE
node, verify service health from an independent vantage, only then the
second. Blue/green with a health gate between hops — always serial.
- **Pi-hole v6: `dnsmasq_lines` in pihole.toml are passed to FTL's embedded
dnsmasq, which rejects some valid-dnsmasq options (e.g. `no-negcache`)
with "bad option" and dnsmasq then never starts — the container looks
"Up" but serves nothing on :53. Use first-class toml settings instead;
never sed-edit pihole.toml without a single-node health-gated rollout.
- **DNS outage recovery path:** workstation resolv.conf dies with the LAN
Pi-holes; use `tailscale status` peer IPs + the chokepoint env overrides
(NETINFRA01_HOST/NETINFRA02_HOST, VM_IP) to reach hosts by Tailscale
while names are unresolvable. Fix one node, verify, then the other.
- **dhcpd + AppArmor:** on the netinfra pair, dhcpd can only read configs
under `/etc/dhcp/**` — staging a candidate config in /tmp fails dhcpd -t
with "Permission denied" even as root. Stage as /etc/dhcp/*.candidate,
test, then move into place.
## Questions (NON-NEGOTIABLE)
**NEVER use a harness "question"/"ask user" tool** (structured prompts,
modal forms, tabbed questions). Banned across every project, every harness.
They are not portable, not version-controlled, and bypass the git record.
**All questions go in the current `questions-v(N).md` file** — write the
question; the human edits the answer inline in the same file. **Version up
the filename each time answers land** (v1 → v2 → v3...): create
`questions-v2.md` with resolved Q&A marked, new questions appended. This
preserves the history of each Q&A round. Synthesize resolved Q&A into
Discourse (decisions) and Redmine (work items). See `BASELINE-PROMPT.md`
§10.
## Documentation policy (IMPORTANT)
**Discourse is the canonical source of truth for all knowledge documentation.**
Knowledge docs (architecture, runbooks, references, audits, policies) have been
migrated to [community.turnsys.com](https://community.turnsys.com/c/vp-techops)
as wiki topics in the **VP TechOps** category.
All `.md` files in this repo (except `AGENTS.md` and `LICENSE`) are now
**pointers** that link to their corresponding Discourse topic. **Do not update
documentation content in git** — edit the Discourse wiki topic instead. Git
edit history no longer serves as the documentation changelog; Discourse
preserves wiki edit history automatically.
Code (scripts, configs, playbooks) still lives in git as the source of truth
for executables. Only *documentation* moved to Discourse.
**Redmine wikis stay disabled — Discourse is the sole documentation space**
(founder ruling, 2026-09-02). All Redmine projects have the wiki module off
(verified 0/62, 2026-09-02); never enable it on new or existing projects.
## Top-level files
All `.md` files now point to Discourse. The key pointers:
| File | Points to | Discourse topic |
|------|-----------|-----------------|
| `README.md` | Project overview | [#296](https://community.turnsys.com/t/296) |
| `STATUS.md` | Ticket index + infra summary | [#297](https://community.turnsys.com/t/297) |
| `docs/docmap.md` | Documentation index | [#296](https://community.turnsys.com/t/296) |
Work tracking stays in [Redmine](https://projects.knownelement.com).
## Repository Layout
```
k8s/ k3s cluster setup scripts (HA control plane over Tailscale) + docs/
proxmox/ Proxmox fleet docs (hardware audit, capacity, storage)
awx/ Ansible AWX deployment (k3s + AWX Operator)
HomeAssistant/ Redirect stub — HA tooling + reference packs moved to KNEL/pfv-bms (tools/, docs/reference-packs/) [#762]
tests/ Test suite + VM validation harness + remote.sh SSH chokepoint
scripts/ Framework: git hooks, rule engine (check-rules.sh), shared lib
docs/ Server-build docs, docmap index, and archive
archive/ Historical/superseded code (provisioning -> replaced by KNELIAC project;
KNELServerBuild -> live patterns now in KNEL/monitoring, remainder
preserved under archive/KNELServerBuild/ [#474])
vendor/ Vendored KNELShellFramework
```
- **2026-09-03 repo split (#769)** — domain tooling moved out of this catch-all
into dedicated KNEL repos (working dirs `~/projects/KNEL/<repo>`; full git
history retained here):
| body of work | repo | ticket |
|---|---|---|
| SIEM/syslog (Wazuh) | KNEL/siem | #335 |
| DNS/NTP/DHCP + switches | KNEL/netinfra | #694 |
| Fleet CA | KNEL/ca | #697 |
| Monitoring/OAM | KNEL/monitoring | #343, #773 |
| Inventory/CMDB | KNEL/inventory | #705, #774 |
| Facilities (PDU/console/UPS/sensors) | KNEL/facilities | #458 |
| Performance/tuning | KNEL/perf | #709, #732, #737 |
| Backups (PBS) | KNEL/backups (new stub) | #315, #684 |
| Outbound mail (PMG) | KNEL/mailgateway (new stub) | #696 |
| Secrets mgmt | KNEL/secrets (new stub) | #770 |
| apt-satellite (deb version control) | KNEL/apt-satellite (new stub) | #771 |
| Security Onion (ops side) | KNEL/security-onion (new stub) | #772 |
Unit tests moved with their code (netinfra, facilities). IaC consumers:
KNEL/KNELIAC. Env map + snapshot lifecycle: Discourse t/331.
- **Server provisioning moved to KNELIAC**: The
[`archive/provisioning/`](archive/provisioning/) tree is historical. Active
server provisioning lives in the **KNELIAC** project at
`/home/reachableceo/projects/KNEL/KNELIAC`.
- **Non-bash files**: Some files under `archive/provisioning/Agents/` have `.sh`
extension but are PHP (shebang `#!/usr/bin/php`). Skip in syntax checks.
- **Remote access (NON-NEGOTIABLE):** ALL SSH/SCP to ANY host MUST go
through the chokepoint scripts — [`tests/remote.sh`](tests/remote.sh)
(Proxmox hosts + all VMs) or
`~/projects/KNEL/netinfra/dns-cluster-setup/remote-dns.sh`
(DNS infra hosts: netinfra-01/02, netboot; lives in KNEL/netinfra [#769]). NEVER call
`ssh`/`scp` directly — the harness blocks raw ssh and the command scanner
rejects it. There are no exceptions.
- **DNS names ONLY (NON-NEGOTIABLE):** NEVER use IP address literals
(neither LAN nor Tailscale IPs) in any command, script, or config.
ALWAYS use DNS names. For Proxmox hosts: `PROX_HOST=<dns-name>`. For VMs:
`VM_IP=<dns-name>`. For conman/SNMP/any tool: pass the DNS name. If a
DNS name does not resolve, fix it in DNS (Technitium) or consult the
[system inventory — Discourse #307](https://community.turnsys.com/t/307).
Do NOT fall back to IP literals. This rule eliminates the per-session
discovery tax of finding the right IP for each host.
- **How to access a production VM:**
1. Look up the DNS name in the [system inventory — Discourse #307](https://community.turnsys.com/t/307).
2. `VM_IP=<dns-name> VM_USER=root bash tests/remote.sh vm '<command>'`
3. If the name does not resolve from the workstation, use `PROX_HOST=<proxmox-node>`
and run `qm guest cmd <vmid> network-get-interfaces` (visibility only —
NOT execution) to find the Tailscale DNS name, then access via that.
## Git Policy
1. **ALWAYS commit + push. Never wait. Ever.** After each logical unit of
work, immediately stage, commit (conventional format), and push to
`origin/main`. Do NOT pause for review. Do NOT ask permission. Do NOT
stop to "let the user read it first." The user reviews rendered markdown
on Gitea AFTER it is pushed. **This rule overrides any default
conservative commit-and-hold behavior** — if a default instruction says
"never commit unless asked," that instruction is superseded here.
2. **Atomic commits.** Each commit coherent on its own.
3. **Conventional format**: `feat(scope): desc`, `fix(scope): desc`,
`docs: desc`, `refactor(scope): desc`, `test(scope): desc`.
4. **All shell scripts MUST pass `shellcheck` before commit.** No exceptions.
Run it via the wrapper:
```bash
bash tests/shellcheck.sh # whole repo
bash tests/shellcheck.sh ups/*.sh # specific files
```
This invokes `koalaman/shellcheck:stable` through Docker (no native binary
needed). Fix every reported finding — including `info`-level — or add a
targeted `# shellcheck disable=SCxxxx # <reason>` directive with a
justification. A script that emits any diagnostic is a protocol violation.
Non-bash scripts (PHP with `.sh` shebang `#!/usr/bin/php`, etc.) are exempt.
## Cross-linking (NON-NEGOTIABLE)
**Every artifact cross-references its related artifacts with CLICKABLE
links.** House rule (Charles, 2026-09-02, [#743]) — becomes hard
requirement once GLPI/CMDB change control is online (ITIL/ITSM ramp):
- Redmine ticket comments link the commits they describe (full Gitea URL,
e.g. `https://git.knownelement.com/KNEL/PFVCluster/commit/<sha>`).
- Commit bodies carry the full URL of the corresponding Redmine comment
(`https://projects.knownelement.com/issues/NNN#note-M`) — the "meat"
lives in the comment; the commit body points at it.
- Discourse posts link both the Redmine ticket and the Gitea repo/commit;
Redmine tickets and repo docs link their Discourse topics.
- One click from ANY system reaches the other two. An artifact without its
links is not done.
- When GLPI is online: change requests become the entry point for infra
changes and get cross-referenced into Redmine/Discourse/Gitea the same
way; CMDB items link their tickets and vice versa.
Mechanical enforcement (rule-engine check for the ticket-comment URL in
commit bodies when `[#NNN]` is present) is tracked under [#441].
## CI/CD (fleet standard, 2026-09-04)
**Every gitea repo runs CI via Gitea Actions.** Runner:
`ukrrs-pfv-gitea-runner` (act_runner 0.2.13, compose at
`~/projects/ukrrs/ci/gitea-runner/` on the pfv workstation) — registered
**org-scoped to KNEL** with label **`ultix` ONLY**; workflows must use
`runs-on: ultix` or they queue forever.
Repo duty: carry `.gitea/workflows/ci.yml` — lint for the repo's language
(pinned container images, never `:latest`) + a secret-scan step, on `dev`
and `release` pushes. Branch model: `dev` = WIP (CI must pass), `release` =
production/deploy branch (CI is the final gate).
Full standard: [Discourse topic #333](https://community.turnsys.com/t/333).
Tracking: [#784](https://projects.knownelement.com/issues/784).
## TDD & Linting
- **Red/green TDD for all code.** Mandatory (founder 2026-08-27). Interim
relaxation applies ONLY until the Ansible/IaC transition (week of 9/1);
from then on every playbook/script ships with its failing test first.
`scripts/test.sh` is the local gate; `tests/validation` + `tests/security`
run on sectestbed targets.
- **Linters on all code, as early as possible.** shellcheck zero-warning
(including info-level) is already enforced pre-commit.
## IaC Codification (shadow tracking) — NON-NEGOTIABLE
**Every manual/direct change to a fleet system must be codified same-day as
an AWX playbook item.** The fleet converges to 100% IaC (founder mandate,
2026-08-27, #454).
- Made a manual change? Add a checklist item to #454 (or a child ticket) in
the same session — what changed, where, exact commands, and any quirks
(e.g. "needed udevadm trigger after NUT install").
- Work is NOT "done" until the manual change is codified or explicitly
ticketed for codification.
- New manual fixes during incidents: fix first, codify immediately after.
- Playbooks live in this repo, tested through the `sectestbed-*` fleet.
## DNS change discipline (NON-NEGOTIABLE)
**Any change to Technitium records/zones (dns-cli or API), dhcpd, or the
live DNS/DHCP configs MUST end with a same-session git SoR sync.**
Founder rule, 2026-09-02 — drift-check exists because git went stale; do
not reopen that gap:
1. Make the change (serial + health-gated across the redundant pair).
2. `cd ~/projects/KNEL/netinfra && bash dns/technitium/zone-snapshot.sh` — refresh the DZ
snapshots from the primary.
3. `cd ~/projects/KNEL/netinfra && bash dns/drift-check.sh` — must print `ALL IN SYNC`.
4. Commit + push the refreshed snapshots in the same session.
## Redmine Tracking Policy
**Redmine is the system of record for all work.** Do not track status,
checklists, or TODOs in repo files. Use Redmine tickets instead.
**Tracker discipline:** OAM/ops/feature/audit tickets use tracker **Support
(3)** — NOT Bug. The redmine-cli `create` defaults to Bug; always correct the
tracker after create (python escape hatch: `tracker_id=3`). [2026-08-27]
- **URL:** https://projects.knownelement.com
- **Version:** Potential to Kinetic Ready (due 2026-09-30)
- **Project:** Known Element Enterprises - Technology & Facility Services (id 55)
### Rules
1. **Every piece of work** (feature, fix, deployment, config change) gets a
Redmine ticket. If one doesn't exist, create it.
2. **Reference tickets in docs and commits** using `[#NNN]` notation.
Example: `[#367] Rebuilt k3s control plane after cnode wipe`.
3. **When work completes**, update the ticket: set done ratio to 100%,
add a note describing what was done and where the code lives.
**NEVER close a ticket without explicit user permission.** You may
SUGGEST a close when the result is clearly scoped and delivered. If
it's ambiguous whether the work is truly complete, don't suggest a
close — leave that decision to the user.
4. **Operations Status** lives on Discourse ([topic #297](https://community.turnsys.com/t/297)) — update that wiki topic if the infrastructure summary needs refreshing. The `STATUS.md` file in git is now a pointer only.
5. **Link code to tickets** — ticket descriptions and notes should reference
the relevant file paths in this repo (e.g., the ups tooling in KNEL/facilities).
### CLI access (read + write)
Tickets are managed via the `redmine-cli` container, invoked directly
with `docker run` (no wrapper script). Full command reference, patterns, and
the subtask escape hatch live in the CLI's own `AGENTS.md`
(`~/projects/ukrrs/connectors/KNEL-AIMiddleware/tooling-cli/redmine/AGENTS.md`) — read it
for anything beyond the basics.
```bash
# Connection sanity check (run first in any session):
docker run --rm --env-file ~/.creds/redmine.env \
git.knownelement.com/reachableceo/redmine-cli:latest whoami
# Your queue (project 55):
docker run --rm --env-file ~/.creds/redmine.env \
git.knownelement.com/reachableceo/redmine-cli:latest list --assigned-to-me -p 55
# Show / create / update / close:
docker run --rm --env-file ~/.creds/redmine.env \
git.knownelement.com/reachableceo/redmine-cli:latest show 367
docker run --rm --env-file ~/.creds/redmine.env \
git.knownelement.com/reachableceo/redmine-cli:latest create -p 55 -s "Subject" -d "desc"
docker run --rm --env-file ~/.creds/redmine.env \
git.knownelement.com/reachableceo/redmine-cli:latest update 367 -n "Done: committed in abc123" --done-ratio 100
docker run --rm --env-file ~/.creds/redmine.env \
git.knownelement.com/reachableceo/redmine-cli:latest close 367
```
Key IDs: project **55** (`technicaloperations`), user **5** (`reachableceo`).
Statuses: New(1), In Progress(2), Resolved(3,closed), Feedback(4), Closed(5),
Rejected(6). New subtasks go to **Feedback (4)**. Tracker **3** = Support.
**Gotcha:** `create` has no `--parent` flag — to make a subtask, use the
`python-redmine` escape hatch inside the container (see
`tooling-cli/redmine/AGENTS.md`). Always `show` a ticket before updating it.
Credentials (`REDMINE_URL`/`REDMINE_API_KEY`) live in the centralized store
at `~/.creds/redmine.env`.
## Documentation Workflow
**Discourse is the source of truth for all knowledge docs.** After any work
session, an agent MUST:
1. **Update Discourse wiki topics** — if infrastructure facts changed (new
VM, IP change, host retired), edit the relevant wiki topic at
[community.turnsys.com/c/vp-techops](https://community.turnsys.com/c/vp-techops).
2. **Update the Operations Status topic** ([#297](https://community.turnsys.com/t/297))
if tickets were opened or closed.
3. **Grep for stale paths in code** — `grep -rn 'old/path' --include='*.sh'`
after any rename or restructure. Fix all references in the same commit.
4. **If a new top-level directory was created**, update:
- `AGENTS.md` → "Repository Layout" code block
- `AGENTS.md` → "Key Scripts" table (if it has an entrypoint script)
- Create a new Discourse wiki topic for any documentation
5. **Self-audit before commit.** Code changes must be internally consistent.
Documentation changes go to Discourse, not git.
### CLI access (read + write)
Wiki topics are managed via the `discourse-cli` container, invoked directly
with `docker run` (no wrapper script). Full command reference, patterns, and
the raw-API escape hatch live in the CLI's own `AGENTS.md`
(`~/projects/ukrrs/connectors/KNEL-AIMiddleware/tooling-cli/discourse/AGENTS.md`) — read it
for anything beyond the basics.
```bash
# Connection sanity check (run first in any session):
docker run --rm --env-file ~/.creds/discourse.env \
git.knownelement.com/reachableceo/discourse-cli:latest whoami
# List VP TechOps topics:
docker run --rm --env-file ~/.creds/discourse.env \
git.knownelement.com/reachableceo/discourse-cli:latest ls -c vp-techops
# Show a topic / edit a wiki post (find post id via `show`):
docker run --rm --env-file ~/.creds/discourse.env \
git.knownelement.com/reachableceo/discourse-cli:latest show 297
docker run --rm --env-file ~/.creds/discourse.env \
git.knownelement.com/reachableceo/discourse-cli:latest update <post_id> -b "new markdown body"
```
VP TechOps = category **74**. Key topics: #296 (project overview), #297
(ops status), #298 (audit log). The API user is trust-level 4 but **not
admin** — admin-only ops (category creation, setting the wiki flag) will
403; surface those to the user rather than retrying. **Gotcha:**
`update`/`delete` take a post **id**, not a post number. Never create a new
topic for an update to existing knowledge — edit the wiki post in place.
Credentials (`DISCOURSE_URL`/`DISCOURSE_API_KEY`/`DISCOURSE_API_USERNAME`)
live in the centralized store at `~/.creds/discourse.env`.
## Key Scripts
| Script | Purpose |
|--------|---------|
| [`scripts/check-rules.sh`](scripts/check-rules.sh) | Rule audit engine (shellcheck, image pinning, Discourse pointers, required files) |
| [`scripts/setup-hooks.sh`](scripts/setup-hooks.sh) | Install git hooks (pre-commit, pre-push) |
| [`scripts/pve-snapshot.sh`](scripts/pve-snapshot.sh) | PVE host snapshot/backup pass — hard-protects prod before risky windows (verified 2026-09) |
| [`scripts/audit-fleet.sh`](scripts/audit-fleet.sh) | Fleet-wide audit sweep (hosts, services, drift) |
| [`scripts/garden.sh`](scripts/garden.sh) | Repo/workspace gardening loop (strays, cruft, staleness) |
| `scripts/kuma-*.py` suite | Uptime Kuma management (add-ping, fleet-sync, inventory, maintenance, notifications, regroup) |
| [`tests/remote.sh`](tests/remote.sh) | **SSH chokepoint** — all Proxmox host + sandbox VM access routes here |
| KNEL/netinfra: `dns-cluster-setup/remote-dns.sh` | SSH chokepoint for DNS infra hosts (netinfra-01/02, netboot) |
| KNEL/netinfra: `dns/technitium/zone-snapshot.sh` | Refresh git SoR Technitium zone snapshots from the primary — **required after every DNS change** |
| KNEL/netinfra: `dns/drift-check.sh` | Live-vs-git drift audit for DHCP/DNS/NTP/zones — must be green before any DNS-adjacent push |
| KNEL/inventory: `seed/from-inventory.sh` | Regenerate the CMDB seed CSV from the Discourse #307 inventory raw (re-run when #307 changes) |
| `redmine-cli` container | Redmine CLI (ticket read/write via `docker run`; see `tooling-cli/redmine/`) |
| `discourse-cli` container | Discourse CLI (wiki topic read/write via `docker run`; see `tooling-cli/discourse/`) |
| `dns-cli` container | Technitium DNS CLI (zones, list, add, delete, search, flush; see `tooling-cli/dns/`) |
| [`tests/vm-validation.sh`](tests/vm-validation.sh) | Deploy + validate on sandbox VM |
| [`tests/run-tests.sh`](tests/run-tests.sh) | Test suite |
| KNEL/netinfra: `dns-cluster-setup/setup.sh` | DNS cluster replication |
| [`k8s/install-cp.sh`](k8s/install-cp.sh) | Bootstrap k3s HA control plane |
| KNEL/facilities: `powerman/setup.sh` | Configure Cyclades PDU via powerman |
| KNEL/facilities: `console/setup.sh` | Configure serial console access via ser2net + conman |
| KNEL/facilities: `ups/setup.sh` | Configure NUT (Network UPS Tools) for UPS monitoring |
| KNEL/perf: `deploy-tuning.sh` | Deploy perf tunings |
| KNEL/perf: `scripts/` | Read-only audit: probe-storage, probe-network, conman-console, snmp-switch-audit, probe-drift, audit-vm-disks, audit-guest-io, deploy-tuned-guests |
## Switch Console Access (conman)
Switch configs are pulled via serial console through a conman + ser2net
stack on pfv-tsys4.
- **conmand server:** pfv-tsys4 (port 7890)
- **ser2net:** pfv-tsys4, TCP ports 2001-2006
(do NOT connect to ser2net directly — it conflicts with conman's
persistent sessions; always use the conman client)
- **Script:** KNEL/perf: `scripts/conman-console.py`
— drives console sessions read-only via PTY. No expect/tcl required.
- **Command files:** KNEL/netinfra: `switches/` — `.cmds` files
with switch-specific show commands
- **Query available consoles:** `conman -d pfv-tsys4 -q`
```bash
# Pull a switch config (example):
CONMAN_SERVER=pfv-tsys4:7890 python3 ~/projects/KNEL/perf/scripts/conman-console.py \
--console pfv-r5-core-01 --cmds ~/projects/KNEL/netinfra/switches/pfv-r5-core-01.cmds
```
| Console name | TCP port | Device |
|--------------|----------|--------|
| pfv-r5-core-01 | 2001 | Dell PowerConnect 5448 (rack 5 core, mgmt+storage) |
| pfv-r3-tor-mgmt-01 | 2002 | Dell PowerConnect 5324 (rack 3 mgmt TOR) |
| pfv-r3-tor-stor-01 | 2003 | Dell PowerConnect 5324 (rack 3 storage TOR) |
| pfv-rrinfra-rtr | 2004 | Cisco router (rrinfra) |
| pfv-r2-tor-01 | 2005 | Rack 2 TOR switch |
| pfv-r6-mgmt-01 | 2006 | Rack 6 management switch |
## Mandatory: Document ALL Infrastructure Changes
**This is non-negotiable. Every infrastructure change (VM config, disk
cache, network setting, service config, storage migration) MUST be
documented BEFORE moving to the next task step — not "later" or "at the
end."**
Required for EVERY infrastructure change:
1. **Redmine ticket** — create one if none exists. Reference as `[#NNN]`.
2. **Discourse audit log** — reply to topic
[#298](https://community.turnsys.com/t/298) with a dated entry (what
changed, why, where).
3. **Discourse relevant wiki topic** — update the architecture/reference
topic if the change affects documented infrastructure facts (storage
#300, network #299, k8s #305, etc.).
4. **Git commit** — if code/config changed in the repo, commit + push
immediately per the Git Policy above.
## Tooling
- `gh`, `docker`, `jq` available on the workstation.
- No native shellcheck — use `bash tests/shellcheck.sh` (Docker wrapper
`koalaman/shellcheck:stable`). ALL scripts must pass including info-level.
- For raw API calls not covered by the Redmine/Discourse CLIs, use
`python3` inside the CLI Docker containers (escape hatch pattern in
`tooling-cli/{discourse,redmine}/AGENTS.md`).
- `curl`/`wget`/`httpie` may be blocked by some harnesses. Use the CLIs or
the python-in-Docker escape hatch for HTTP writes.
## Key Commands Quick Reference
```bash
# Tests + validation:
bash tests/run-tests.sh # test suite
bash tests/vm-validation.sh # VM validation
bash tests/shellcheck.sh # shellcheck whole repo
bash tests/shellcheck.sh path/to/*.sh # shellcheck specific files
# Redmine + Discourse sanity checks (run first in any session):
docker run --rm --env-file ~/.creds/redmine.env \
git.knownelement.com/reachableceo/redmine-cli:latest whoami
docker run --rm --env-file ~/.creds/discourse.env \
git.knownelement.com/reachableceo/discourse-cli:latest whoami
# Proxmox host access (DNS names only — never IPs):
PROX_HOST=pfv-tsys5 bash tests/remote.sh prox 'qm list'
PROX_HOST=pfv-tsys5 bash tests/remote.sh prox 'pvesm status'
# Production VM access (DNS names only — never IPs):
VM_IP=tsys-librenms VM_USER=root bash tests/remote.sh vm 'systemctl status cron'
# DNS infra access:
bash ~/projects/KNEL/netinfra/dns-cluster-setup/remote-dns.sh netinfra01-root 'systemctl status docker'
# Switch console (read-only config pull — DNS names only):
CONMAN_SERVER=pfv-tsys4:7890 python3 ~/projects/KNEL/perf/scripts/conman-console.py \
--console pfv-r5-core-01 --cmds ~/projects/KNEL/netinfra/switches/pfv-r5-core-01.cmds
```
## Credential Management
- API keys currently in `.env` files under KNEL-AIMiddleware (gitignored).
- User goal: migrate all keys to Hashicorp Vault
(`vault.knownelement.com`). No vault token present yet
(`~/.vault-token` missing). Track as high-priority TODO.
## Key Docs
→ **All documentation lives on Discourse:**
[community.turnsys.com/c/vp-techops](https://community.turnsys.com/c/vp-techops)
→ All `.md` files in this repo are pointers to Discourse topics.
→ **Complete Linux System Inventory:**
[Topic #307](https://community.turnsys.com/t/307) — every Linux system
(hosts + VMs + physical), with Tailscale IPs, DNS names, SSH access
status, and tuned profiles. Reference this for monitoring coverage,
access management, and hostname consistency.
## Project Context
Solo-founder R&D Proxmox cluster in a private residence. Shoestring budget.
Production lives on a Cloudron VPS in Reston VA. See the
[Operations Status topic](https://community.turnsys.com/t/297) for the ticket
index and infrastructure summary.
All work is tracked in [Redmine](https://projects.knownelement.com)
(version: Potential to Kinetic Ready, due 2026-09-30).
- **HA 2026.8 snmp platform traps:** temperature/current-classed sensors with
`state_class: measurement` + non-numeric render ('unknown') are dropped AT
ADD-TIME (ValueError) when a host is asleep — entity vanishes entirely. Fix:
no state_class (stats live in VM) + templates render EMPTY on no-match.
- **Legacy snmp extend OID encoding:** suffix = strlen + ASCII decimal per
char ("garagepdu" = 9.103.97.114.97.103.101.112.100.117 — count carefully).
- **Old APC PDUs (AOS ~2.7):** SNMPv1 only; subtree WALKS die silently, small
single-table walks + GETs work. Garage AP7830 current OID:
.1.3.6.1.4.1.318.1.1.12.2.3.1.1.2.1 (tenths of amps). Relay pattern via
tsys-librenms (the allowed poller vantage) for v1-only/ACL-locked gear.
- **Debian snmpd on monitoring VMs:** snmpd may fail to bind 127.0.0.1:161
(exit 1, "Error opening specified endpoint") when something squats
localhost — bind only the LAN/TS addresses in agentaddress.
- **iDRAC6/7 SNMP:** agents exist but hw MIB tables unpopulated on old fw;
IPMI SDR via iDRAC SSH (`racadm`, one command per exec, legacy syntax on
iDRAC6: `racadm config -g cfgOobSnmp -o ...`) is the data path.
- **jq in one-liner state checks:** use `.state` — not `.s` — a wrong key
silently returns null and looks like an entity outage (cost an hour tonight).
- **HA config flows are REST** (`POST /api/config/config_entries/flow`),
not websocket; websocket only monitors. ha-ws-call.py + REST flow are the
two levers for agent-driven HA setup.
- **mvdan/sh `source a b` silently ignores file b** — source env files
one per line.
- **HA todo lists (founder HUD):** items are hidden from `/api/states` —
fetch via `POST /api/services/todo/get_items?return_response=true`.
Renaming works ONLY with a minimal `todo/update_item` payload
(`entity_id`, `item`=<uid>, `rename`); passing `status`/`due` back earns
a 400. HUD summaries carry `[#NNN]` Redmine prefixes — sync completed
items back to Redmine when the founder asks.
- **HA package YAML: NEVER append a second top-level `sensor:` key** — YAML
duplicate keys are last-one-wins and silently shadow the whole first block
at config load (#344 outage 2026-09-02: all 22 fleet temps down after the
restart that pulled the file). Merge new sensors into the existing block.
- **HA 2026.8 snmp add-time crash, general form:** ANY unit-bearing sensor
with a non-numeric native_value at ADD (host down renders '' / literal
string "unknown"; value_template can never yield None — even `{{ none }}`
is stringified to "None") = ValueError, entity vanishes till manual
reload. `check_config` does NOT catch it. Durable pattern: unit-less raw
snmp transport + `template:` layer carrying unit/device_class under the
ORIGINAL unique_ids (KNEL/pfv-bms packages/plant_snmp.yaml header).
- **HA entity_id comes from the NAME slug, not unique_id** — if a sensor's
name and unique_id diverge, cross-references must use the name slug
(garage PDU amps lesson: semantic read a nonexistent `..._amps_raw`).
- **HA entity_id collisions spawn `_2` entities; registry remove no-ops for
LIVE entities** (they re-register on their next write). Clean `_2` by
jq-editing `.storage/core.entity_registry` with core STOPPED, backup
first. Orphaned (dead) entries remove fine via websocket
(ha-ws-call.py — pass commands as JSON files; shell quoting eats argv).
- **git_pull add-on: `ha addons start` is a no-op while it's running** —
use `restart` to force an immediate pull. It auto-restarts core for
files in its restart list; otherwise pull ≠ reload (config lands on
disk, running core keeps the old one until a platform reload/restart).
- **Restore drills / heavy I/O NEVER on hosts running prod VMs** (2026-09-03
incident, #684): qmrestore of a 32G disk to scratch VMID 5599 on tsys1
starved VM 100 (pfv-bms) I/O ~40 min, degraded prod HA. Drills run dev-lane
only (tsys5); prod restores are server-throttled via PBS traffic-control.
- **`pkill -f <pat>` over the ssh chokepoint self-matches** your own remote
command line and kills the session (exit 255, no output). Use the bracket
trick: `pkill -f "[q]mrestore"`.
- **pvesh backup jobs take `-vmid`** (comma list), not `-vmids`; aborted
`qmrestore` leaves a create-locked stub VM — `qm unlock <id>` then destroy.
- **PBS 4.x gotchas:** prune retention lives in `prune-job` (datastore keep-*
opts are gone); API tokens need their OWN acl entry — user-only ACL yields
"Cannot find datastore" from PVE.
- **PVE bridge port + hotplug NIC (#780):** a re-registered NIC (USB move
changes its enumeration path) comes back UP but is NOT re-enslaved to its
bridge — `ifup <nic>` is not enough; run `ifup <bridge>` and confirm
`bridge link` shows `state forwarding`. NFS `hard` mounts mostly self-heal,
but `nconnect=4` mounts can stay wedged on stale channels after a flap:
force-remount (`umount -f -l`, recreate the mountpoint dir if the lazy
umount removed it, fresh mount). `pvesm status` "active" can be stale —
verify with a real read, not the status line.
- **PMG pmgproxy restart hang (#696):** `systemctl restart pmgproxy` can hang
in "deactivating (stop)" on old workers ignoring TERM. If stuck >90s:
`systemctl kill -s KILL pmgproxy && systemctl start pmgproxy` — API/UI
only, mail flow unaffected. Postfix restarts don't hang.
+141
View File
@@ -0,0 +1,141 @@
# Global Baseline Prompt
> The canonical set of working principles for every agent (AI or human) across
> every project. This is the source of truth; project `AGENTS.md` files inherit
> and specialize it. Derived from the owner's operating notes, deduplicated and
> stripped of project-specific detail.
>
> **Phase context:** we have exited the "move fast and loose" phase. This is
> production infrastructure — it is in production right now. The bar is the bar.
---
## 1. You are an employee, not a lone genius
- **Stop over-thinking.** Get to code and output faster. Explore with code; gather ground truth. Do not burn tokens reasoning about things a quick command can answer.
- **Ask questions early.** Use a git-tracked `questions-v(N).md` file (see §9) that the human reviews inline. Questions, answers, and the reasoning behind decisions are often more important than the code. Capture and synthesize them to Discourse/Redmine.
- **Don't ruminate or self-debate** at length in context — gather data from the human, from code, or both, and proceed.
- **You are not alone.** Ask for guidance when you need it.
## 2. Token efficiency is a hard constraint
- The owner has a quota. Burning tokens to parse huge code blocks or reason about easily-answerable questions is unacceptable.
- **Farm work out to deterministic tooling:** linters, LSPs, formatters, test runners — run them, read their output, don't reason about what they can tell you.
- If an LSP is wired up for the language (code or docs), use it. If not, pull a Docker image and lint inside it. Prose linting belongs in a Docker image too.
- Use `STATUS.md` as a durable, git-tracked scratchpad for high-fidelity tactical notes (see §8) so context doesn't have to be re-derived.
- Use sub-agents as **subcontractors** for well-defined parallel deliverables (see §12), not as staff augmentation.
## 3. Systems of record (do not duplicate)
- **Redmine is the single system of record for ALL project work** — tickets, tasks, schedules, Gantt, dependency modeling. Use the `redmine-cli` tool. Gitea issues are not used.
- **Discourse is the single system of record for documentation.** It is all Markdown. Use the `discourse-cli` tool. Do not author long-form docs in gitea.
- **Git-tracked `.md` files should be stubs** that point to the relevant Discourse URL. Short operational files that must live next to code (e.g. `AGENTS.md`, `STATUS.md`) are the exception.
- Engineering already works this way. Operations now does too.
- **Cross-referencing is mandatory.** Every Redmine ticket links to its Discourse doc; every Discourse doc links to its Redmine ticket(s); every commit references `[#NNN]`; every PR links to both Redmine and Discourse. Keep them in sync at all times.
- **Code, docs, and tests must be kept in sync at all times.** When you change code, update the corresponding docs (Discourse) and tests in the same commit. Never leave them out of sync.
## 4. Infrastructure change approval workflow
For infrastructure (operations) work, agents do NOT execute changes
without explicit human approval. The workflow is:
1. **Prepare:** Agent creates scripts, configs, and a plan. Documents
the exact changes, blast radius, and rollback procedure in a
Redmine ticket.
2. **Review:** Human reviews and approves (or rejects) via Redmine.
3. **Execute:** Agent applies the change ONLY after approval.
4. **Verify:** Agent verifies the change worked and documents results
in Discourse.
This applies to all production infrastructure: network configs, host
tuning, VM settings, DNS records, switch configs, etc. Read-only
audits and probes do NOT require approval — only changes that alter
system behavior.
Code projects (software development) follow normal git/PR workflow and
do not require per-change Redmine approval.
## 5. Git workflow
- **Use the `tea` CLI for pull requests.**
- **Work smart off master** generally. Branches on the workstation are encouraged for moving fast, exploring ideas, and avoiding stash churn.
- **Once work leaves the workstation, it goes through a PR.**
- Branching strategy is open to per-project discussion.
- **Commits and PRs must cross-link** to Redmine tickets (`[#NNN]` in subject or body) and Discourse docs (in PR body).
## 6. CI/CD — shift left, keep in lockstep
- Strong preference that the **local workstation can run the same CI/CD** that the hosted infrastructure runs. Maintain them in lockstep across all projects.
- **The further left CI/CD runs, the better.** Catch it before push, before PR, before merge.
- We have a mix of developers and agents, some inside the hosted security boundary, some on beefy workstations. CI/CD must work for all of them.
## 7. Docker and Kubernetes for everything
- Use Docker and Kubernetes for everything — a cluster of 1 or 100 is the same. Don't presume scale. Containers are containers; k8s is k8s.
- **All development work happens in containers** — custom, off-the-shelf, or a mix. `docker pull` freely without asking.
- **Container naming: never use Docker's default.** Always name with a project prefix (e.g. `<project>-<service>`).
- Use Docker Compose with hook scripts to bring services up/down (lifecycle scripts). See `~/projects` for established examples.
- **Host hygiene is inviolable.** The host runs only `git`, `docker`, and standard Unix utilities (`awk`, `sed`, `grep`, `cut`, `tr`, `jq`, `find`, `xargs`, etc.). No language runtimes or package-managed tools beyond the base OS.
- **One-off utility needed?** `docker pull` a pinned image and run it ephemerally. Do not install on the host.
- **Broadly useful tool?** Create a Redmine ticket requesting the human add it via the system package manager. Do not install it yourself.
## 8. Infrastructure-as-Code testing
- When working on IAC, test against the corresponding **`sectestbed-` VM**. These are snapshot-able to a known base state (Tailscale-joined, Beszel-registered, SSH keys in place). The base state evolves; the delta of tested code shrinks over time.
- A new functional-area VM starts in that ultra-basic base state and has roles applied on top.
- **`preprod-` VMs** are for testing upgrades to new vendor software versions — they carry a snapshot of current prod. Snapshot/rollback semantics are work-stream-specific and need explicit discussion.
- Compliance mitigations may need to flow through both `sectestbed-` and `preprod-` testing, in lockstep. Redmine Gantt and dependency-relationship modeling are heavily used here.
- **Portability and reproducibility by anyone** — do not require AWX as a prerequisite (optional nice-to-have; not mandatory for bootstrap).
## 9. STATUS.md — scratchpad, not system of record
- **STATUS.md is a durable, git-tracked scratchpad for token efficiency.** It is not the system of record (Redmine is).
- The agent fully owns STATUS.md; the human only consumes it.
- Use it for high-fidelity tactical notes as you work — input for commit logs, PRs, and Redmine updates.
- The harness todo tool is fine for tracking *current* work; STATUS.md is the durable cross-session record.
- **STATUS.md has an Inbox section.** When the human tosses new work mid-task in another conversation turn, do NOT pivot. Log it in the Inbox. If it's materially different, spin up a Redmine ticket.
## 10. Questions file — `questions-v(N).md`
- **NEVER use a harness "question"/"ask user" tool** (structured prompts,
modal forms, tabbed questions, etc.). Ever. They are banned across every
project. They are not portable across harnesses, they don't version
history, and they bypass the git-tracked record. This is non-negotiable.
- **All questions go in the git-tracked `questions-v(N).md` file.** Write
the question; the human edits the answer inline in the same file. This
preserves history, works under every harness, and keeps reasoning next
to the answer. Version up (`questions-v2.md`, …) when a round lands.
- Capture questions in a versioned file: `questions-v1.md`, `questions-v2.md`, …
- The human reviews and edits it inline. Version it when a round of answers goes in.
- Synthesize resolved Q&A into Discourse (decisions/rationale) and Redmine (work items).
## 11. Belt-and-suspenders protocol enforcement
- Enforce the rules in **two layers**: prose policy in `AGENTS.md` and mechanical enforcement in git `pre-commit` / `pre-push` hooks. Harness-specific hooks (e.g. Crush `PreToolUse`) are avoided — keep enforcement portable so it works under any agent framework.
- Never rely on memory or prose alone.
## 12. Gardening — keep docs from sprawling
- Run a routine **gardening loop.** Agents are disciplined with code but tend to sprawl `.md` files everywhere.
- Keep docs, code, and tests in sync at all times.
- `scripts/garden.sh` reports `.md` sprawl and files that should be migrated to Discourse.
## 13. Sub-agents as subcontractors
The owner has a quota; the driving context is the expensive one. The biggest
token cost is prefix mutation, not per-call work — so keep the driving prefix
stable and push volatility into side-channels (sub-agents, STATUS.md).
- **Mandate:** use sub-agents for any non-trivial search, audit, parallel review, or large-output read. **Never read 10+ files sequentially** — batch them into 2-3 agent calls.
- **Self vs. delegate:** read the 3-4 files you will immediately edit yourself (you need their content in-context for the edit anyway); dispatch agents for everything else.
- **Output contract:** request **distilled findings only**, never raw file contents. Specify the output format in the prompt. A sub-agent that returns a 500-line file dump has failed the contract.
- **Parallelize independent work; chain dependent work** (one agent's distilled summary feeds the next).
- **Why:** keeps the main context lean and preserves the cached prompt prefix.
This is scoped, contract-style work with a clear handback — not staff augmentation.
## 14. TDD and linting
- **Red/green TDD for all code.** Write the failing test first.
- **Linters on all code, as early as possible.** Be token-efficient — let deterministic tools find the issues.
- **Prefer Unix utilities (awk, sed, grep, cut, tr, etc.) for file editing and text processing** over built-in harness edit tools when feasible. Harness edit tools can be unreliable with whitespace/indentation; Unix tools are deterministic and exact.
-3
View File
@@ -1,3 +0,0 @@
#Global Variables used by the framework
export ProjectIncludes="1"
+15
View File
@@ -0,0 +1,15 @@
# HomeAssistant/ — moved to KNEL/pfv-bms
Everything Home-Assistant-related now lives in the config source of
truth: [KNEL/pfv-bms](https://git.knownelement.com/KNEL/pfv-bms)
(public, AGPLv3).
| Moved | New home |
|---|---|
| `ha-ws-call.py` (websocket CLI) | `tools/ha-ws-call.py` |
| `packs/*.yaml` (reference packs) | `docs/reference-packs/` (reference only, not loaded) |
Operational facts (host access, REST API env, deploy pipeline) are
documented in the pfv-bms README and on
[Discourse VP TechOps](https://community.turnsys.com/c/vp-techops).
This stub remains for link compatibility. [#762]
+50
View File
@@ -0,0 +1,50 @@
# Makefile — convenience dispatch to scripts/.
#
# Not required. The scripts in scripts/ are the real entry points and work
# standalone. This file just gives you short verbs if you're at a terminal.
#
# In Mode 2 (Hermes/OWUI/MCP), agents call the scripts directly or via API —
# they don't need this file.
# Project-specific overrides for check-rules.sh
export PROJECT_DOC_EXEMPT ?= legacy-* AGENTS.md STATUS.md WORKING.md README.md ADOPTING.md LICENSE .env.example questions-v1.md BASELINE-PROMPT.md PATTERNS.md
export PROJECT_DISCOURSE_HOST ?= community.turnsys.com
.PHONY: setup validate fast lint test garden up down status clean help
help: ## Show available targets
@grep -E '^[a-zA-Z_-]+:.*?## ' $(MAKEFILE_LIST) | awk 'BEGIN{FS=":.*?## "}{printf " \033[36m%-12s\033[0m %s\n", $$1, $$2}'
setup: ## Install git hooks
@bash scripts/setup-hooks.sh
validate: ## Full rule audit (includes tests)
@bash scripts/check-rules.sh
fast: ## Fast rule audit (pre-commit equivalent)
@bash scripts/check-rules.sh --fast
lint: ## Lint shell scripts (shellcheck via docker)
@docker run --rm -v "$$(pwd):/mnt" koalaman/shellcheck:stable \
$$(find . -path ./.git -prune -o -path ./.tmp -prune -o -path ./vendor -prune -o -path ./node_modules -prune -o \( -name '*.sh' -o -name '*.bash' \) -print | sed 's|^\./|/mnt/|') || true
test: ## Run the test suite (override per project)
@bash scripts/test.sh
garden: ## Doc-sprawl / Discourse-migration report
@bash scripts/garden.sh
up: ## Bring up the docker-compose stack
@bash scripts/up.sh
down: ## Bring down the docker-compose stack
@bash scripts/down.sh
status: ## Show repo status snapshot
@echo "== branch =="; git branch --show-current 2>/dev/null || echo "(no branch)"
@echo "== last commit =="; git log --oneline -1 2>/dev/null || true
@echo "== working tree =="; git status --short 2>/dev/null || echo "(not a git repo)"
@echo "== STATUS.md head =="; sed -n '1,12p' STATUS.md 2>/dev/null || echo "(no STATUS.md)"
clean: ## Remove build/test artifacts (override per project)
@echo "make clean: nothing to clean — override this in your project's Makefile."
+247
View File
@@ -0,0 +1,247 @@
# Cross-Project Pattern Extraction
**Scope:** 18 projects surveyed across two machines:
- **Local (`ultix-streaming`):** 12 infra/stack projects — `dotfiles`, `EngineeringWorkstation`, `EngStack`, `football`, `hermes-rceo-streaming`, `KNEL-AIMiddleware`, `KNELIAC`, `netbird`, `PFVCluster`, `TSYSDevStack-SupportStack-Cloudron`, `WorkstationStack`, `TSYSDevStack-SupportStack-LocalWorkstation`.
- **Remote (`ultix-offstage`):** 6 personal/business projects — `CharlesNWybleResume`, `rceo-automation-espanso-private`, `RCEO-PersonalAssistant`, `ReachableCEO-Profile-FullTimeEmployment`, `RevGen-TimeForMoney`, `ThreeYearPlan`.
This document records the patterns that recur, the inconsistencies between them,
and the decisions the repo makes to drive inconsistency to zero.
---
## 1. AGENTS.md — recurring structure
A canonical section set emerges from the 12 files that have one (8 local + 4 remote):
| Section | Prevalence |
|---|---|
| Git commit & push policy | 12/12 |
| Project overview / context | 11/12 |
| Key commands / build cheat sheet | 10/12 |
| Repository layout (ASCII tree) | 9/12 |
| Status / journal maintenance | 9/12 |
| Conventions & naming | 8/12 |
| DO / DON'T | 7/12 |
| Validation / testing | 7/12 |
| Key files reference | 6/12 |
| Quick start / onboarding | 5/12 |
**Decision:** the template's `AGENTS.md` ships all ten sections in a fixed order,
so every project has the same skeleton. Filler is bracketed for replacement.
---
## 2. Shared cross-project preferences (the house style)
These recur strongly and are baked into the template:
1. **Auto-commit + push is mandatory and non-negotiable.** Stated with ALL-CAPS force in nearly every file. The template states the override-once and lets the hooks enforce it.
2. **Atomic commits — one logical change per commit.** Universal.
3. **Conventional-commit format** (`feat:`/`fix:`/`docs:`/…). Universal, but the **💘 Crush attribution footer is inconsistent**: mandated only by `football` and `KNELIAC`; absent on the remote entirely. **Decision:** the template mandates the footer once, in the AGENTS.md commit block.
4. **Docker-only host hygiene.** The single most consistent convention across *both* machines (`KNEL-AIMiddleware`, `EngStack`, `LocalWorkstation`, `RCEO-PersonalAssistant`, `ThreeYearPlan`, `CharlesNWybleResume`). **Decision:** the template's `check-rules.sh` enforces host hygiene via the Docker-only rule (no `:latest` tags, container naming) and `scripts/docker-run.sh` provides the canonical container wrapper. The Crush-specific `enforce-bash.sh` hook was removed for harness portability; the policy lives in AGENTS.md prose.
5. **Sub-agents encouraged.** Explicit in `football`, `ThreeYearPlan`, `RCEO-PersonalAssistant`. **Decision:** standard section in the template.
6. **No secrets in git (infra genre).** Inverted on the remote private repos (`RCEO-PersonalAssistant` deliberately commits credentials). **Decision:** the template keeps the infra stance (`.env` gitignored) as the default; private repos may opt out.
**Inconsistencies the template resolves:**
- `shellcheck` is mandated by `PFVCluster` and `football` but unmentioned elsewhere → the template mandates it for all shell, via Docker.
- "Keep the host clean" is phrased four different ways → one canonical phrasing.
- Auto-commit override language varies in strength → stated once, authoritatively.
---
## 3. Git hooks — the biggest gap, now standardized
**Finding:** custom git hooks are essentially unused. Only **3 of 18** projects have any:
| Project | Hook | Mechanism | Installer |
|---|---|---|---|
| `football` | `pre-commit` (4-check SDLC gate: shellcheck, unit tests, coverage, doc-sync) | `core.hooksPath githooks/` | `scripts/setup-githooks.sh` |
| `KNEL-AIMiddleware` | `pre-push` (block on dirty tree) | `core.hooksPath .githooks/` | **none** (manual `git config`) |
| `RCEO-PersonalAssistant` | `pre-commit` (fast audit) + `pre-push` (full Docker tests) | **copy** into `.git/hooks/` | `scripts/setup-hooks.sh` |
Three different install mechanisms; one project has no installer at all.
**Decision:** the template uses the **copy** approach (most portable: works on any clone, no config mutation, idempotent) with a single `scripts/setup-hooks.sh`, and combines both policy philosophies:
- `pre-commit` = fast audit with a **hot-path bypass** for `STATUS.md` / `JOURNAL.md` / `WORKING.md` (so frequent status commits stay frictionless) — proven in RCEO.
- `pre-push` = full audit (incl. `scripts/test.sh`) — proven in RCEO — plus the dirty-tree gate — proven in KNEL-AIMiddleware.
---
## 4. Crush hooks — an entirely untapped capability (except one project)
**Finding:** Crush `hooks` are used by **exactly one project**`RCEO-PersonalAssistant`. Every other `crush.json` (and most projects have none) defines only `lsp`/`mcp`/`options`, never `hooks`. All process enforcement elsewhere is prose in AGENTS.md — manually enforced, violable.
RCEO's crush.json wires five `PreToolUse` hooks that form a behavioral guardrail layer complementing the git hooks:
| Hook | Matcher | Effect |
|---|---|---|
| `block-todos.sh` | `^todos$` | bans the todos tool; WORKING.md is the only tracker |
| `enforce-bash.sh` | `^bash$` | blocks banned commands + host language toolchains |
| `enforce-rules.sh` | `^(edit\|write\|multiedit)$` | blocks banned file types; TDD reminder on source edits |
| `audit-before-git.sh` | `^bash$` | runs the fast audit before any `git commit`/`git push` |
| `exit-protocol.sh` | `.*` | blocks stopping while WORKING.md has unchecked tasks |
**Decision:** the Crush hooks were studied as the proof-of-concept and their *policies* (ban todos, Docker-only, banned file types, audit before git, exit protocol) were ported into AGENTS.md prose and `check-rules.sh` mechanical checks. The harness-specific `crush.json` + `hooks/` layer was **deliberately removed** for portability — the user is shifting away from Crush to OpenWebUI/Hermes, so all enforcement is git hooks + prose only.
---
## 5. The rules engine — `check-rules.sh`
RCO's `check-rules.sh` is the best reusable artifact found. Its core abstractions:
- A `check()` accumulator classifying each rule as **pass / warn / fail**, with `FAIL > 0 ⇒ exit 1` and `WARN` non-fatal.
- A `--fast` mode (used by pre-commit) that silences per-check output and skips slow checks.
- Self-checking categories: shellcheck, Docker image pinning (no `:latest`), required-files manifest, doc freshness (STATUS.md/JOURNAL.md touched today), git state, hooks-installed, WORKING.md completion, unresolved `CNW:` markers.
**Decision:** the template's `check-rules.sh` generalizes this — language-specific checks (Go test coverage, gofmt) are dropped or made opt-in via `PROJECT_*` env; universal checks (shellcheck, image pinning, required files, doc freshness, hygiene, WORKING.md, CNW markers, merge-conflict markers) are kept. The `check()` accumulator and `--fast`/`--quiet` flags move into `scripts/lib/common.sh` so any script can reuse them.
---
## 6. Scripts — driving the shebang/boilerplate chaos to zero
**Findings:**
- **Three incompatible shebang variants** across projects: `#!/usr/bin/env bash` (best), `#!/bin/bash`, `#!/usr/bin/bash` (non-portable).
- **Four different `set`-flag policies**: `set -euo pipefail` (best), `set -e` only, `set -uo pipefail`, none. KNEL-AIMiddleware uses `set -e` only — piped failures silently swallowed; `BuildAll.sh`'s `docker compose … | tail` can report green on failure.
- **The ANSI color block is copy-pasted into ~10 scripts.**
- **`log_*` helpers are redefined per project** (and per-script in KNEL-AIMiddleware, where they aren't even functions).
- **Repo-root resolution** is re-rolled in nearly every script.
- **No shared cross-project library exists.** Only `EngStack` factors helpers (`scripts/lib.sh` build-time + `lib/common.sh` runtime).
**Decision:** `scripts/lib/common.sh` provides the deduplicated primitives once: ANSI colors, `log_info/ok/warn/error/step`, `have()`, `die()`, `repo_root()`, `as_root()`, `docker_run()`, and the `init_counters`/`check`/`print_summary_and_exit` audit helpers. All template scripts use `#!/usr/bin/env bash` + `set -euo pipefail`.
---
## 7. Task runners — five conventions, standardized to one
**Findings:** no `Makefile`/`Taskfile`/`justfile` exists anywhere. Five incompatible conventions for the same semantic verbs:
| Project | Convention |
|---|---|
| football | `./run.sh <subcommand>` |
| KNEL-AIMiddleware | `scripts/<Verb>.sh` |
| EngStack | `scripts/<verb>-<noun>.sh` |
| PFVCluster | `tests/<name>.sh` |
| hermes-agent | npm workspace scripts |
The recurring semantic targets — **build, test, lint, validate, status, clean, setup** — map cleanly onto one target set that doesn't exist.
**Decision:** a single `Makefile` provides `setup`, `validate`, `fast`, `lint`, `test`, `status`, `clean`, `help`. It is pure dispatch to `scripts/`; projects override `test`/`clean` for their stack. CI, hooks, and humans now share one set of verbs.
---
## 8. CI — present in exactly one project
Only `hermes-agent` has CI (22 GitHub Actions workflows — an exemplar: change-detection orchestrator, reusable-workflow lanes, SHA-pinned actions, supply-chain/OSV scans, live PR-comment bot). The other 17 have none.
**Decision:** the template doesn't ship CI (it's stack-dependent), but `scripts/check-rules.sh` + `scripts/garden.sh` give any future workflow a uniform entry point. The hermes-agent `ci.yml` orchestrator is the documented growth path.
---
## 9. Crush configuration layers — documented but unused
Crush defines a priority chain (`$HOME/.config/crush/crush.json``crush.json``.crush.json`). In practice:
- The **global** config (`dotfiles`) is trivial (attribution style only) — no shared LSP/MCP baseline.
- Only **2 projects** have a project-local `crush.json` (KNEL-AIMiddleware: 4 LSP + 36 MCP via wrappers; RCEO-PersonalAssistant: hooks only).
- The highest-priority `.crush.json` layer is **unused anywhere**.
- **No project sets `model` or `provider`** in config.
**Decision:** no Crush config is shipped. The hooks layer was removed for harness portability. LSP/MCP remain project-local concerns.
---
## 10. Crush memory — a singleton, now a template
Only `PFVCluster` has operational memory (`.crush/memory/operational.md`). Its format — Identity → Tracking Systems (with READY/PARTIAL/BLOCKED status) → Tooling → Access chokepoints → Key commands → Mandatory rules → Credential TODOs — is strong and reusable.
**Decision:** the template ships `.crush/memory/operational.md` as a fill-in skeleton (optional, for projects using Crush).
---
## 11. KNEL-AIMiddleware wrapper ecosystem — a standardization opportunity (not in the template)
`KNEL-AIMiddleware` hand-maintains **38 near-identical** wrapper scripts (`mcp-*-wrapper.sh`, `lsp-*-wrapper.sh`) for its MCP/LSP fleet. ~90% of each file is duplicated boilerplate (container-cleanup stanza, `docker run -i --rm --name`, env passthrough). Drift is already visible (default style, double-passed creds, one structurally-different LSP wrapper). There is **no generator**; they were hand-cloned by an agent over many sessions.
**Recommendation (out of scope for the template):** replace the 38 files with one parameterized launcher driven by a declarative `servers.yaml` manifest, plus a `gen-wrappers.sh`. The template's `docker-run.sh` + `lib/common.sh` show the direction; the same idea applies at fleet scale.
---
## 12. Workflow documents — WORKING.md / STATUS.md / JOURNAL.md
The remote personal genre contributes a discipline absent from infra: a **task/state/document triad** enforced by hooks:
- **WORKING.md** — the only task tracker (todos tool banned); commits blocked while `- [ ]` remain; cleared before responding.
- **STATUS.md** — agent-maintained dashboard humans read; staleness is a warning.
- **JOURNAL.md** — append-only ADR/pattern log; no today-entry is a warning.
- **`CNW:` markers** — flag unresolved questions for the human; empty markers are a warning.
**Decision:** all four are in the template, and check-rules.sh enforces their freshness/completion.
---
## Standardization scorecard
| Inconsistency | Was | Now |
|---|---|---|
| Shebangs | 3 variants | `#!/usr/bin/env bash` |
| `set` flags | 4 policies | `set -euo pipefail` |
| Color/log boilerplate | copy-pasted ~10× | `lib/common.sh` |
| Task runner | 5 conventions | one `Makefile` |
| Git hook install | 3 mechanisms, 1 missing | one `setup-hooks.sh` (copy) |
| Git hook coverage | 3/18 projects | every project, two hooks |
| Crush hooks | 1/18 projects | not used (harness-agnostic; git hooks only) |
| Rules engine | 1 project (Go-specific) | generalized, project-agnostic |
| Crush memory | 1 project | skeleton in template (optional, Crush-only) |
| AGENTS.md structure | bespoke per project | canonical 10-section skeleton |
| 💘 commit footer | 2/18 projects | mandated in template (harness-agnostic) |
---
## 13. Global baseline prompt integration
After the initial extraction, the owner shared a draft "global baseline prompt"
(an Apple Note of operating observations gathered over weeks of working with
agents). It encodes principles that cut across every project and that the
template now reflects. The cleaned canonical version lives in
[`BASELINE-PROMPT.md`](BASELINE-PROMPT.md); the project-level specialization is
in `AGENTS.md`. New artifacts and checks added:
### New policy encoded
| Principle (baseline §) | How the template enforces it |
|---|---|
| **Stop over-thinking; ask early** (§1, §9) | `questions-v1.md` skeleton + AGENTS.md "Questions" section; required-files check fails if absent |
| **Token efficiency / farm to tooling** (§2) | `scripts/check-rules.sh` runs shellcheck in Docker; AGENTS.md "Working Style" forbids parsing huge code in context |
| **Redmine = SoR for work; Discourse = SoR for docs** (§3) | AGENTS.md "Systems of Record" section; `scripts/garden.sh` flags oversized non-Discourse `.md` |
| **Git .md = stubs to Discourse** (§3) | `scripts/garden.sh` reports oversized `.md` lacking a Discourse URL |
| **`tea` CLI for PRs; off-workstation → PR** (§4) | AGENTS.md "Git Workflow" policy |
| **Shift-left CI/CD, lockstep local + hosted** (§5) | AGENTS.md "CI/CD" section; `scripts/check-rules.sh --fast` runs at pre-commit |
| **Docker/k8s for everything; container naming** (§6) | new container-naming rule in `check-rules.sh` (every compose service needs `container_name`); `docker-compose.yml.example` + `scripts/up.sh` / `scripts/down.sh` |
| **STATUS.md = scratchpad, not SoR; has Inbox** (§8) | STATUS.md reframed; Inbox section for mid-task interruptions ("don't pivot") |
| **Questions file** (§9) | `questions-v1.md` artifact + required-files check |
| **Belt-and-suspenders enforcement** (§10) | already present (git hooks + Crush hooks) — now documented as policy |
| **Gardening loop** (§11) | `scripts/garden.sh` |
| **Sub-agents as subcontractors** (§12) | AGENTS.md "Working Style" |
| **TDD + linters** (§13) | AGENTS.md "TDD & Linting"; `enforce-rules.sh` TDD reminder (opt-in via `PROJECT_SOURCE_SUFFIXES`) |
### New template artifacts
- `BASELINE-PROMPT.md` — the cleaned canonical global prompt (lives in the meta root; projects inherit).
- `questions-v1.md` — git-tracked question log skeleton.
- `docker-compose.yml.example` — lifecycle template with project-prefix naming.
- `scripts/up.sh`, `scripts/down.sh` — compose lifecycle wrappers.
- `scripts/garden.sh` — doc-sprawl / Discourse-migration report.
### New `check-rules.sh` checks
- **Container naming** (§6): every service in a `docker-compose*.y*ml` must declare `container_name:`; failure otherwise.
- **Questions file required**: `questions-v1.md` must exist (required-files manifest).
### New Makefile targets
`scripts/garden.sh`, `scripts/up.sh`, `scripts/down.sh` — standard scripts across every project.
### What stayed project-level (not globalized)
- `sectestbed-` / `preprod-` VM testing semantics — IAC-workstream-specific; left as AGENTS.md prose, not a mechanical check.
- Redmine Gantt / dependency modeling — workflow-specific.
- Specific DNS/cloudron/SITER production details — out of scope.
@@ -1,6 +0,0 @@
option rfc3442-classless-static-routes code 121 = array of unsigned integer 8;
send host-name = gethostname();
request subnet-mask, broadcast-address, time-offset, routers,
domain-name, host-name,
rfc3442-classless-static-routes;
@@ -1 +0,0 @@
install cramfs /bin/true
@@ -1 +0,0 @@
install dccp /bin/true
@@ -1 +0,0 @@
install freevxfs /bin/true
@@ -1 +0,0 @@
install hfs /bin/true
@@ -1 +0,0 @@
install hfsplus /bin/true
@@ -1 +0,0 @@
install jffs2 /bin/true
@@ -1 +0,0 @@
install rds /bin/true
@@ -1 +0,0 @@
install sctp /bin/true
@@ -1 +0,0 @@
install squashfs /bin/true
@@ -1 +0,0 @@
install tipc /bin/true
@@ -1 +0,0 @@
install udf /bin/true
@@ -1 +0,0 @@
install usb-storage /bin/true
-7
View File
@@ -1,7 +0,0 @@
driftfile /var/lib/ntp/ntp.drift
leapfile /usr/share/zoneinfo/leap-seconds.list
server pfv-netboot.knel.net
restrict 127.0.0.1
restrict ::1
interface ignore wildcard
interface listen 127.0.0.1
@@ -1 +0,0 @@
/.*/ tsysrootaccount@knel.net
-34
View File
@@ -1,34 +0,0 @@
#!/bin/bash
#curl -s http://dl.turnsys.net/omsa.sh|/bin/bash
gpg --keyserver hkp://pool.sks-keyservers.net:80 --recv-key 1285491434D8786F
gpg -a --export 1285491434D8786F | apt-key add -
echo "deb http://linux.dell.com/repo/community/openmanage/930/bionic bionic main" > /etc/apt/sources.list.d/linux.dell.com.sources.list
wget http://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman-curl-client-transport1_2.6.5-0ubuntu3_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman-client4_2.6.5-0ubuntu3_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman1_2.6.5-0ubuntu3_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman-server1_2.6.5-0ubuntu3_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/universe/s/sblim-sfcc/libcimcclient0_2.2.8-0ubuntu2_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/openwsman_2.6.5-0ubuntu3_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/multiverse/c/cim-schema/cim-schema_2.48.0-0ubuntu1_all.deb
wget http://archive.ubuntu.com/ubuntu/pool/universe/s/sblim-sfc-common/libsfcutil0_1.0.1-0ubuntu4_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/multiverse/s/sblim-sfcb/sfcb_1.4.9-0ubuntu5_amd64.deb
wget http://archive.ubuntu.com/ubuntu/pool/universe/s/sblim-cmpi-devel/libcmpicppimpl0_2.0.3-0ubuntu2_amd64.deb
dpkg -i libwsman-curl-client-transport1_2.6.5-0ubuntu3_amd64.deb
dpkg -i libwsman-client4_2.6.5-0ubuntu3_amd64.deb
dpkg -i libwsman1_2.6.5-0ubuntu3_amd64.deb
dpkg -i libwsman-server1_2.6.5-0ubuntu3_amd64.deb
dpkg -i libcimcclient0_2.2.8-0ubuntu2_amd64.deb
dpkg -i openwsman_2.6.5-0ubuntu3_amd64.deb
dpkg -i cim-schema_2.48.0-0ubuntu1_all.deb
dpkg -i libsfcutil0_1.0.1-0ubuntu4_amd64.deb
dpkg -i sfcb_1.4.9-0ubuntu5_amd64.deb
dpkg -i libcmpicppimpl0_2.0.3-0ubuntu2_amd64.deb
apt update
apt -y install srvadmin-all
touch /opt/dell/srvadmin/lib64/openmanage/IGNORE_GENERATION
#logout,login, then run
# srvadmin-services.sh enable && srvadmin-services.sh start
-57
View File
@@ -1,57 +0,0 @@
#Boilerplate and support functions
FrameworkIncludeFiles="$(ls -1 --color=none ../../../Framework-Includes/*)"
IFS=$'\n\t'
for file in ${FrameworkIncludeFiles[@]}; do
source "$file"
done
unset IFS
ProjectIncludeFiles="$(ls -1 --color=none ../../../Project-Includes/*)"
IFS=$'\n\t'
for file in ${ProjectIncludeFiles[@]}; do
source "$file"
done
unset IFS
print_info "Setting up librenms agent..."
cat ../../Agents/librenms/distro > /usr/local/bin/distro
chmod +x /usr/local/bin/distro
if [ ! -d /usr/lib/check_mk_agent ]; then
mkdir -p /usr/lib/check_mk_agent
fi
if [ ! -d /usr/lib/check_mk_agent/plugins ]; then
mkdir -p /usr/lib/check_mk_agent/plugins
fi
if [ ! -d /usr/lib/check_mk_agent/local ]; then
mkdir -p /usr/lib/check_mk_agent/local
fi
cat ../../Agents/librenms/check_mk_agent > /usr/bin/check_mk_agent
chmod +x /usr/bin/check_mk_agent
cat ../../Agents/librenms/check_mk@.service > /etc/systemd/system/check_mk@.service
cat ../../Agents/librenms/check_mk.socket > /etc/systemd/system/check_mk.socket
systemctl enable check_mk.socket
systemctl start check_mk.socket
#Modules commented out below, we will roll out on systems that use them, most of the fleet doesn't use those modules
cat ../../Agents/librenms/dmi.sh > /usr/lib/check_mk_agent/local/dmi.sh
cat ../../Agents/librenms/dpkg.sh > /usr/lib/check_mk_agent/local/dpkg.sh
#cat ../../Agents/librenms/mysql.sh > /usr/lib/check_mk_agent/local/mysql.sh
cat ../../Agents/librenms/ntp-client > /usr/lib/check_mk_agent/local/ntp-client
#cat ../../Agents/librenms/ntp-server.sh > /usr/lib/check_mk_agent/local/ntp-server.sh
cat ../../Agents/librenms/os-updates.sh > /usr/lib/check_mk_agent/local/os-updates.sh
cat ../../Agents/librenms/postfixdetailed > /usr/lib/check_mk_agent/local/postfixdetailed
cat ../../Agents/librenms/postfix-queues > /usr/lib/check_mk_agent/local/postfix-queues
#cat ../../Agents/librenms/smart.sh > /usr/lib/check_mk_agent/local/smart
#cat ../../Agents/librenms/smart.sh.config > /usr/lib/check_mk_agent/local/smart.config
chmod +x /usr/lib/check_mk_agent/local/*
@@ -1,10 +0,0 @@
#!/bin/bash
#secharden-2fa
#Coming very soon, 2fa for webmin/cockpit/ssh
#libpam-google-authenticator
#https://www.ogselfhosting.com/index.php/2024/03/21/enabling-2fa-for-cockpit/
#https://webmin.com/docs/modules/webmin-configuration/#two-factor-authentication
#https://www.digitalocean.com/community/tutorials/how-to-set-up-multi-factor-authentication-for-ssh-on-ubuntu-18-04
@@ -1,52 +0,0 @@
#!/bin/bash
set -o errexit
set -o nounset
set -o pipefail
set -o functrace
export PS4='(${BASH_SOURCE}:${LINENO}): - [${SHLVL},${BASH_SUBSHELL},$?] $ '
function error_out()
{
echo "Bailing out. See above for reason...."
exit 1
}
function handle_failure() {
local lineno=$1
local fn=$2
local exitstatus=$3
local msg=$4
local lineno_fns=${0% 0}
if [[ "$lineno_fns" != "-1" ]] ; then
lineno="${lineno} ${lineno_fns}"
fi
echo "${BASH_SOURCE[0]}: Function: ${fn} Line Number : [${lineno}] Failed with status ${exitstatus}: $msg"
}
trap 'handle_failure "${BASH_LINENO[*]}" "$LINENO" "${FUNCNAME[*]:-script}" "$?" "$BASH_COMMAND"' ERR
export DL_ROOT
DL_ROOT="https://dl.knownelement.com/KNEL/FetchApply/"
# Material herein Sourced from
# https://cisofy.com/documentation/lynis/
# https://jbcsec.com/configure-linux-ssh/
# https://opensource.com/article/20/5/linux-security-lynis
# https://forum.greenbone.net/t/ssh-authentication/13536
# openvas
#lynis
#Auditd
curl --silent ${DL_ROOT}/ConfigFiles/AudidD/auditd.conf > /etc/audit/auditd.conf
# Systemd
curl --silent ${DL_ROOT}/ConfigFiles/Systemd/journald.conf > /etc/systemd/journald.conf
# logrotate
curl --silent ${DL_ROOT}/ConfigFiles/Logrotate/logrotate.conf > /etc/logrotate.conf
@@ -1,108 +0,0 @@
#!/bin/bash
#Framework variables are read from hee
source $PROJECT_ROOT_PATH/Framework-ConfigFiles/FrameworkVars
#Boilerplate and support functions
for framework_include_file in ../Framework-Includes/*; do
source "$framework_include_file"
done
for project_include_file in ../Project-Includes/*; do
source "$project_include_file"
done
# Actual script logic starts here
# Sourced from
# https://complianceascode.readthedocs.io/en/latest/manual/developer/01_introduction.html
# https://github.com/ComplianceAsCode/content
# https://github.com/ComplianceAsCode
#apparmor
#enforcing
#enabled in bootloader config
#aide
#auditd
#disable auto mounting
#disable usb storage
#motd
#remote login warning banner
#Ensure time sync is working
#systemd-timesync
#ntp
#chrony
#password complexity
#password expiration warning
#password expiration time
#password hashing algo
#fix grub perms
if [ "$IS_RASPI" = 0 ] ; then
chown root:root /boot/grub/grub.cfg
chmod og-rwx /boot/grub/grub.cfg
chmod 0400 /boot/grub/grub.cfg
fi
#disable auto mounting
systemctl --now disable autofs || true
apt-get -y --purge remove autofs || true
#disable usb storage
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/usb_storage.conf > /etc/modprobe.d/usb_storage.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/dccp.conf > /etc/modprobe.d/dccp.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/rds.conf > /etc/modprobe.d/rds.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/sctp.conf > /etc/modprobe.d/sctp.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/tipc.conf > /etc/modprobe.d/tipc.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/cramfs.conf > /etc/modprobe.d/cramfs.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/freevxfs.conf > /etc/modprobe.d/freevxfs.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/hfs.conf > /etc/modprobe.d/hfs.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/hfsplus.conf > /etc/modprobe.d/hfsplus.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/jffs2.conf > /etc/modprobe.d/jffs2.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/squashfs.conf > /etc/modprobe.d/squashfs.conf
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/ModProbe/udf.conf > /etc/modprobe.d/udf.conf
#banners
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/BANNERS/issue > /etc/issue
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/BANNERS/issue.net > /etc/issue.net
curl --silent ${DL_ROOT}/ProjectCode/ConfigFiles/BANNERS/motd > /etc/motd
#Cron perms
if [ -f /etc/cron.deny ]; then
rm /etc/cron.deny || true
fi
touch /etc/cron.allow
chmod g-wx,o-rwx /etc/cron.allow
chown root:root /etc/cron.allow
chmod og-rwx /etc/crontab
chmod og-rwx /etc/cron.hourly/
chmod og-rwx /etc/cron.daily/
chmod og-rwx /etc/cron.weekly/
chmod og-rwx /etc/cron.monthly/
chown root:root /etc/cron.d/
chmod og-rwx /etc/cron.d/
# At perms
rm -f /etc/at.deny || true
touch /etc/at.allow
chmod g-wx,o-rwx /etc/at.allow
chown root:root /etc/at.allow
@@ -1,74 +0,0 @@
#!/bin/bash
for framework_include_file in ../../../Framework-Includes/*; do
source "$framework_include_file"
done
for project_include_file in ../../../Project-Includes/*; do
source "$project_include_file"
done
export SUBODEV_CHECK
SUBODEV_CHECK="$(getent passwd | grep -c subodev || true)"
export LOCALUSER_CHECK
LOCALUSER_CHECK="$(getent passwd | grep -c localuser || true)"
export ROOT_SSH_DIR
ROOT_SSH_DIR="/root/.ssh"
export LOCALUSER_SSH_DIR
LOCALUSER_SSH_DIR="/home/localuser/.ssh"
export SUBODEV_SSH_DIR
SUBODEV_SSH_DIR="/home/subodev/.ssh"
if [ ! -d $ROOT_SSH_DIR ]; then
mkdir /root/.ssh/
fi
curl --silent "${DL_ROOT}"/ProjectCode/ConfigFiles/SSH/AuthorizedKeys/root-ssh-authorized-keys >/root/.ssh/authorized_keys
chmod 400 /root/.ssh/authorized_keys
chown root: /root/.ssh/authorized_keys
if [ "$LOCALUSER_CHECK" -gt 0 ]; then
if [ ! -d $LOCALUSER_SSH_DIR ]; then
mkdir -p /home/localuser/.ssh/
fi
curl --silent "${DL_ROOT}"/ProjectCode/ConfigFiles/SSH/AuthorizedKeys/localuser-ssh-authorized-keys >/home/localuser/.ssh/authorized_keys &&
chown localuser /home/localuser/.ssh/authorized_keys &&
chmod 400 /home/localuser/.ssh/authorized_keys
fi
if [ "$SUBODEV_CHECK" = 1 ]; then
if [ ! -d $SUBODEV_SSH_DIR ]; then
mkdir /home/subodev/.ssh/
fi
curl --silent "${DL_ROOT}"/ProjectCode/ConfigFiles/SSH/AuthorizedKeys/localuser-ssh-authorized-keys >/home/subodev/.ssh/authorized_keys &&
chmod 400 /home/subodev/.ssh/authorized_keys &&
chown subodev: /home/subodev/.ssh/authorized_keys
fi
cat ../../ConfigFiles/SSH/Configs/tsys-sshd-config >/etc/ssh/sshd_config
#Don't deploy this config to a ubuntu server, it breaks openssh server. Works on kali/debian.
export UBUNTU_CHECK
UBUNTU_CHECK="$(distro | grep -c Ubuntu||true)"
if [ "$UBUNTU_CHECK" -ne 1 ]; then
cat ../../ConfigFiles/SSH/Configs/ssh-audit-hardening.conf >/etc/ssh/sshd_config.d/ssh-audit_hardening.conf
chmod og-rwx /etc/ssh/sshd_config.d/*
fi
# Perms on sshd_config
chmod og-rwx /etc/ssh/sshd_config
#todo
# only strong MAC algos are used
@@ -1,27 +0,0 @@
#!/bin/bash
# We don't want to run this on the wazuh server, otherwise bad things happen...
export TSYS_NSM_CHECK
TSYS_NSM_CHECK="$(hostname |grep -c tsys-nsm ||true)"
if [ "$TSYS_NSM_CHECK" -eq 0 ]; then
if [ -f /usr/share/keyrings/wazuh.gpg ]; then
rm -f /usr/share/keyrings/wazuh.gpg
fi
curl -s https://packages.wazuh.com/key/GPG-KEY-WAZUH | gpg --no-default-keyring --keyring gnupg-ring:/usr/share/keyrings/wazuh.gpg --import
chmod 644 /usr/share/keyrings/wazuh.gpg
echo "deb [signed-by=/usr/share/keyrings/wazuh.gpg] https://packages.wazuh.com/4.x/apt/ stable main" > /etc/apt/sources.list.d/wazuh.list
apt-get update
WAZUH_MANAGER="tsys-nsm.knel.net" apt-get -y install wazuh-agent
systemctl daemon-reload
systemctl enable wazuh-agent
systemctl start wazuh-agent
echo "wazuh-agent hold" | dpkg --set-selections
fi
-9
View File
@@ -1,9 +0,0 @@
#!/bin/bash
rm -f /etc/apt/sources.list.d/*
echo "deb http://download.proxmox.com/debian/pve bookworm pve-no-subscription" > /etc/apt/sources.list.d/pve-install-repo.list
wget http://download.proxmox.com/debian/proxmox-release-bookworm.gpg -O /etc/apt/trusted.gpg.d/proxmox-release-bookworm.gpg
apt update && apt -y full-upgrade
apt-get -y install ifupdown2 ipmitool ethtool net-tools lshw
#curl -s http://dl.turnsys.net/newSrv.sh|/bin/bash
+9 -18
View File
@@ -1,19 +1,10 @@
# KNEL FetchApply # README.md
## Repo Issues > **Documentation moved to Discourse — the canonical source of truth.**
https://projects.knownelement.com/project/reachableceo-vptechnicaloperations/timeline >
> **Project overview, architecture, quick start**
## Repo Discussion >
https://community.turnsys.com/c/chieftechnologyandproductofficer/26 > **Read it here:** https://community.turnsys.com/t/296
>
> *Migrated 2026-08-06. This file is kept as a pointer for git-browsing context.
## Repo discription > Do not update content here — edit the Discourse wiki topic instead.*
Known Element Enterprises (the entity serving as the TSYS Group management company) (through its executive leader, the COO) provides core IT/back office systems/services/support on a hands off/fully delegated authortity basis to the CCO and the orgs/members.
One of those functions is the provisoning of Linux servers. This repository is the Infrastructure As Code (IAC) repository for TSYS.
In the future it will be used via FetchApply https://github.com/P5vc/fetch-apply
It is invoked via
curl https://dl.knownelement.com/KNEL/FetchApply/SetupNewSystem.sh |/bin/bash
+10
View File
@@ -0,0 +1,10 @@
# STATUS.md
> **Documentation moved to Discourse — the canonical source of truth.**
>
> **Ticket index + infrastructure summary**
>
> **Read it here:** https://community.turnsys.com/t/297
>
> *Migrated 2026-08-06. This file is kept as a pointer for git-browsing context.
> Do not update content here — edit the Discourse wiki topic instead.*
+67
View File
@@ -0,0 +1,67 @@
#!/usr/bin/bash
# access-matrix.sh — definitive access verification across all Linux Tailscale nodes.
# Uses the correct SSH user(s) per system type, checks sudo where applicable.
# Routes through remote.sh (the only allowed ssh path).
set -u
cd /home/reachableceo/projects/PFVCluster || exit 1
# Policy-excluded systems (never attempt access)
EXCLUDE=':tsys-cloudron:pfv-bms:tsys-umbrel:tsys-ucs-01:tsys-ucs-02:stlpc-bizoffice:ultix-highside:'
# Determine the SSH user(s) for a given hostname and whether sudo is expected.
# Returns "user1:user2:...:sudoflag" where sudoflag is "yes" or "no".
users_for() {
local name="$1"
case "$name" in
pfv-tsys[0-9]) echo "root:no" ;;
*-proxmox-datacenter) echo "root:no" ;;
*-proxmox-pve) echo "root:no" ;;
*-proxmox-pbs) echo "root:no" ;;
*-proxmox-mailgw*) echo "root:no" ;;
*-proxmox-backup*) echo "root:no" ;;
tsys-ucs-*) echo "root:no" ;;
ultix-streaming) echo "root:no" ;;
stlpc-*) echo "root:labuser:no" ;;
ultix-field) echo "ultixfield:yes" ;;
subopi*) echo "subodev:yes" ;;
*) echo "localuser:yes" ;;
esac
}
check_user() {
local ip="$1" user="$2" expect_sudo="$3"
local out sudo
out=$(VM_IP="$ip" VM_USER="$user" bash tests/remote.sh vm 'echo SSHOK; id -un' </dev/null 2>&1 | grep -oE 'SSHOK|keyboard-interactive|Connection refused' | head -1)
case "$out" in
SSHOK)
if [ "$expect_sudo" = "yes" ]; then
sudo=$(VM_IP="$ip" VM_USER="$user" bash tests/remote.sh vm 'sudo -n true 2>/dev/null && echo SUDOOK || echo SUDONO' </dev/null 2>&1 | grep -oE 'SUDOOK|SUDONO' | head -1)
printf '%s(%s)' "$user" "${sudo:-?}"
else
printf '%s(ok)' "$user"
fi
;;
keyboard-interactive) printf '%s(2FA)' "$user" ;;
'Connection refused') printf '%s(NOSSH)' "$user" ;;
*) printf '%s(NOKEY)' "$user" ;;
esac
}
printf '%-32s %-16s %s\n' "NAME" "TS-IP" "ACCESS"
printf '%-32s %-16s %s\n' "----" "-----" "------"
tailscale status 2>/dev/null | awk '$4=="linux" {print $2, $1}' | sort | while read -r name ip; do
[ -n "$name" ] || continue
case "$EXCLUDE" in *":$name:"*) printf '%-32s %-16s %s\n' "$name" "$ip" "EXCLUDED"; continue;; esac
map=$(users_for "$name")
expect_sudo="${map##*:}"
users="${map%:*}"
result=""
IFS=':' read -ra user_list <<< "$users"
for u in "${user_list[@]}"; do
r=$(check_user "$ip" "$u" "$expect_sudo")
[ -z "$result" ] && result="$r" || result="$result $r"
done
printf '%-32s %-16s %s\n' "$name" "$ip" "$result"
done
+60
View File
@@ -0,0 +1,60 @@
#!/bin/sh
# agent-bootstrap.sh
#
# Run INSIDE a guest (via SSH, console, or guest-agent) to bring the
# system fully under agent management in one shot:
# 1. install + enable qemu-guest-agent (VMs only, skipped on bare metal)
# 2. push the agent SSH key to root + AGENT_USER (+ labuser if present)
# 3. grant AGENT_USER passwordless sudo
#
# AGENT_USER defaults to "localuser". Override for systems with a different
# unprivileged agent user:
# AGENT_USER=subodev bash agent-bootstrap.sh
#
# After this runs once, the agent has SSH+sudo immediately.
#
# Usage (from a root shell in the guest):
# bash agent-bootstrap.sh
# AGENT_USER=subodev bash agent-bootstrap.sh
set -eu
KEY='ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIIWms/uCXnjjo4KyxHBcYI2TDHe8OZ2wle6W/0hSRQLu reachableceo@ultix-streaming'
AGENT_USER="${AGENT_USER:-localuser}"
# 1. guest-agent (skip on bare metal — no virtio-serial device)
if command -v systemd-detect-virt >/dev/null 2>&1 && \
[ "$(systemd-detect-virt --vm 2>/dev/null || echo none)" != "none" ]; then
if ! command -v qemu-ga >/dev/null 2>&1; then
if command -v apt-get >/dev/null 2>&1; then
apt-get update
DEBIAN_FRONTEND=noninteractive apt-get install -y qemu-guest-agent
elif command -v dnf >/dev/null 2>&1; then
dnf install -y qemu-guest-agent
elif command -v yum >/dev/null 2>&1; then
yum install -y qemu-guest-agent
else
echo "WARN: no supported package manager; skipping agent install" >&2
fi
fi
systemctl enable --now qemu-guest-agent 2>/dev/null || \
systemctl enable --now qemu-ga 2>/dev/null || true
fi
# 2. SSH key for root + AGENT_USER + labuser (if present)
for u in root "$AGENT_USER" labuser; do
if ! getent passwd "$u" >/dev/null 2>&1; then continue; fi
H=$(getent passwd "$u" | cut -d: -f6)
mkdir -p "$H/.ssh"; chmod 700 "$H/.ssh"
AK="$H/.ssh/authorized_keys"; touch "$AK"; chmod 600 "$AK"
grep -qF "$KEY" "$AK" || echo "$KEY" >> "$AK"
chown -R "$u": "$H/.ssh"
done
# 3. passwordless sudo for AGENT_USER only
if getent passwd "$AGENT_USER" >/dev/null 2>&1 && [ -d /etc/sudoers.d ]; then
echo "${AGENT_USER} ALL=(ALL) NOPASSWD:ALL" > /etc/sudoers.d/010-agent
chmod 440 /etc/sudoers.d/010-agent
fi
echo BOOTSTRAP-DONE
+13
View File
@@ -0,0 +1,13 @@
# LOGFILENAME artifacts: the framework (Logging.sh + PrettyPrint.sh) appends
# every print_info/print_error line to LOGFILENAME, defined as
# "$0.<Weekday>-YYYY-MM-DD-HH:MM:SS.$$". Running any script that sources the
# framework therefore drops a timestamped log file next to it. Ignore these
# everywhere in the repo.
*.Monday-*
*.Tuesday-*
*.Wednesday-*
*.Thursday-*
*.Friday-*
*.Saturday-*
*.Sunday-*
dns-cluster-setup/.export/
+76
View File
@@ -0,0 +1,76 @@
# Agent Guidelines
## Repository Layout
Knowing where things live prevents broken edits:
- **Vendored framework**: `KNELShellFramework` lives at
`vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework/`, **not** at
the repo root. Its includes are under `Framework-Includes/` there. Never
assume `./Framework-Includes` exists relative to the repo root.
- **Self-locating scripts**: All provisioning scripts derive their own
location via `BASH_SOURCE` and compute `PROJECT_ROOT_PATH` from it. They must
**never** depend on the current working directory or on `cd`/`realpath ..`
chains. Run them from anywhere — `sudo bash SetupNewSystem.sh` works.
- **Local config files are the source of truth**: Configs in
`ProjectCode/ConfigFiles/` are read with `cat`/`cp`. Do **not** re-introduce
`curl ${DL_ROOT}/...` downloads from `dl.knownelement.com` — that CDN is
deprecated for this repo.
- **Path variables**: Scripts export `PROJECT_ROOT_PATH`, `CONFIGFILES_PATH`,
`MODULES_PATH`, `SCRIPTS_PATH`, and `AGENTS_PATH` for locating repo content.
- **Non-bash agents**: Some files under `ProjectCode/Agents/` carry a `.sh`
extension but are PHP (e.g. `mysql.sh`, shebang `#!/usr/bin/php`). Syntax
checkers must skip these.
## Git Commit Requirements
When making changes to this repository, ALWAYS:
1. **Commit atomically**: Each logical change should be its own commit
2. **Use conventional commit format**:
- `feat(scope): description` - New feature
- `fix(scope): description` - Bug fix
- `docs: description` - Documentation changes
- `refactor(scope): description` - Code refactoring
- `test(scope): description` - Test additions/changes
- `chore: description` - Maintenance tasks
3. **Write verbose, beautifully formatted messages**:
- Title line (50 chars max)
- Blank line
- Body explaining WHAT and WHY (not how)
- Reference related files/issues
- Include footer with attribution
## Example Commit
```
feat(security-hardening): implement SCAP-STIG compliance logic
Refactor apply script to implement comprehensive security hardening:
- Add GRUB bootloader permission hardening (root:root, mode 0400)
- Disable and remove autofs service per STIG requirements
- Deploy modprobe configurations for kernel module blacklisting
- Create STIG-compliant network protocol blacklist
This ensures servers meet DoD security requirements for production
deployment.
🤖 Generated with [Crush](https://github.com/charmassociates/crush)
Assisted-by: GLM-5 via Crush <crush@charm.land>
```
## Autonomous Git Workflow
**Agents are authorized to commit AND push autonomously. Do not wait to be
asked.** After each logical unit of work:
1. Stage only the files belonging to that logical change.
2. Commit with a conventional, well-formed message (see above).
3. Push to `origin` (`git push`). The branch tracks `origin/main`.
4. Repeat per logical unit.
Group changes so each commit is coherent on its own (a reader should
understand the commit without seeing the others). Never batch unrelated
changes into one commit.
+235
View File
@@ -0,0 +1,235 @@
GNU AFFERO GENERAL PUBLIC LICENSE
Version 3, 19 November 2007
Copyright (C) 2007 Free Software Foundation, Inc. <http://fsf.org/>
Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed.
Preamble
The GNU Affero General Public License is a free, copyleft license for software and other kinds of works, specifically designed to ensure cooperation with the community in the case of network server software.
The licenses for most software and other practical works are designed to take away your freedom to share and change the works. By contrast, our General Public Licenses are intended to guarantee your freedom to share and change all versions of a program--to make sure it remains free software for all its users.
When we speak of free software, we are referring to freedom, not price. Our General Public Licenses are designed to make sure that you have the freedom to distribute copies of free software (and charge for them if you wish), that you receive source code or can get it if you want it, that you can change the software or use pieces of it in new free programs, and that you know you can do these things.
Developers that use our General Public Licenses protect your rights with two steps: (1) assert copyright on the software, and (2) offer you this License which gives you legal permission to copy, distribute and/or modify the software.
A secondary benefit of defending all users' freedom is that improvements made in alternate versions of the program, if they receive widespread use, become available for other developers to incorporate. Many developers of free software are heartened and encouraged by the resulting cooperation. However, in the case of software used on network servers, this result may fail to come about. The GNU General Public License permits making a modified version and letting the public access it on a server without ever releasing its source code to the public.
The GNU Affero General Public License is designed specifically to ensure that, in such cases, the modified source code becomes available to the community. It requires the operator of a network server to provide the source code of the modified version running there to the users of that server. Therefore, public use of a modified version, on a publicly accessible server, gives the public access to the source code of the modified version.
An older license, called the Affero General Public License and published by Affero, was designed to accomplish similar goals. This is a different license, not a version of the Affero GPL, but Affero has released a new version of the Affero GPL which permits relicensing under this license.
The precise terms and conditions for copying, distribution and modification follow.
TERMS AND CONDITIONS
0. Definitions.
"This License" refers to version 3 of the GNU Affero General Public License.
"Copyright" also means copyright-like laws that apply to other kinds of works, such as semiconductor masks.
"The Program" refers to any copyrightable work licensed under this License. Each licensee is addressed as "you". "Licensees" and "recipients" may be individuals or organizations.
To "modify" a work means to copy from or adapt all or part of the work in a fashion requiring copyright permission, other than the making of an exact copy. The resulting work is called a "modified version" of the earlier work or a work "based on" the earlier work.
A "covered work" means either the unmodified Program or a work based on the Program.
To "propagate" a work means to do anything with it that, without permission, would make you directly or secondarily liable for infringement under applicable copyright law, except executing it on a computer or modifying a private copy. Propagation includes copying, distribution (with or without modification), making available to the public, and in some countries other activities as well.
To "convey" a work means any kind of propagation that enables other parties to make or receive copies. Mere interaction with a user through a computer network, with no transfer of a copy, is not conveying.
An interactive user interface displays "Appropriate Legal Notices" to the extent that it includes a convenient and prominently visible feature that (1) displays an appropriate copyright notice, and (2) tells the user that there is no warranty for the work (except to the extent that warranties are provided), that licensees may convey the work under this License, and how to view a copy of this License. If the interface presents a list of user commands or options, such as a menu, a prominent item in the list meets this criterion.
1. Source Code.
The "source code" for a work means the preferred form of the work for making modifications to it. "Object code" means any non-source form of a work.
A "Standard Interface" means an interface that either is an official standard defined by a recognized standards body, or, in the case of interfaces specified for a particular programming language, one that is widely used among developers working in that language.
The "System Libraries" of an executable work include anything, other than the work as a whole, that (a) is included in the normal form of packaging a Major Component, but which is not part of that Major Component, and (b) serves only to enable use of the work with that Major Component, or to implement a Standard Interface for which an implementation is available to the public in source code form. A "Major Component", in this context, means a major essential component (kernel, window system, and so on) of the specific operating system (if any) on which the executable work runs, or a compiler used to produce the work, or an object code interpreter used to run it.
The "Corresponding Source" for a work in object code form means all the source code needed to generate, install, and (for an executable work) run the object code and to modify the work, including scripts to control those activities. However, it does not include the work's System Libraries, or general-purpose tools or generally available free programs which are used unmodified in performing those activities but which are not part of the work. For example, Corresponding Source includes interface definition files associated with source files for the work, and the source code for shared libraries and dynamically linked subprograms that the work is specifically designed to require, such as by intimate data communication or control flow between those
subprograms and other parts of the work.
The Corresponding Source need not include anything that users can regenerate automatically from other parts of the Corresponding Source.
The Corresponding Source for a work in source code form is that same work.
2. Basic Permissions.
All rights granted under this License are granted for the term of copyright on the Program, and are irrevocable provided the stated conditions are met. This License explicitly affirms your unlimited permission to run the unmodified Program. The output from running a covered work is covered by this License only if the output, given its content, constitutes a covered work. This License acknowledges your rights of fair use or other equivalent, as provided by copyright law.
You may make, run and propagate covered works that you do not convey, without conditions so long as your license otherwise remains in force. You may convey covered works to others for the sole purpose of having them make modifications exclusively for you, or provide you with facilities for running those works, provided that you comply with the terms of this License in conveying all material for which you do not control copyright. Those thus making or running the covered works for you must do so exclusively on your behalf, under your direction and control, on terms that prohibit them from making any copies of your copyrighted material outside their relationship with you.
Conveying under any other circumstances is permitted solely under the conditions stated below. Sublicensing is not allowed; section 10 makes it unnecessary.
3. Protecting Users' Legal Rights From Anti-Circumvention Law.
No covered work shall be deemed part of an effective technological measure under any applicable law fulfilling obligations under article 11 of the WIPO copyright treaty adopted on 20 December 1996, or similar laws prohibiting or restricting circumvention of such measures.
When you convey a covered work, you waive any legal power to forbid circumvention of technological measures to the extent such circumvention is effected by exercising rights under this License with respect to the covered work, and you disclaim any intention to limit operation or modification of the work as a means of enforcing, against the work's users, your or third parties' legal rights to forbid circumvention of technological measures.
4. Conveying Verbatim Copies.
You may convey verbatim copies of the Program's source code as you receive it, in any medium, provided that you conspicuously and appropriately publish on each copy an appropriate copyright notice; keep intact all notices stating that this License and any non-permissive terms added in accord with section 7 apply to the code; keep intact all notices of the absence of any warranty; and give all recipients a copy of this License along with the Program.
You may charge any price or no price for each copy that you convey, and you may offer support or warranty protection for a fee.
5. Conveying Modified Source Versions.
You may convey a work based on the Program, or the modifications to produce it from the Program, in the form of source code under the terms of section 4, provided that you also meet all of these conditions:
a) The work must carry prominent notices stating that you modified it, and giving a relevant date.
b) The work must carry prominent notices stating that it is released under this License and any conditions added under section 7. This requirement modifies the requirement in section 4 to "keep intact all notices".
c) You must license the entire work, as a whole, under this License to anyone who comes into possession of a copy. This License will therefore apply, along with any applicable section 7 additional terms, to the whole of the work, and all its parts, regardless of how they are packaged. This License gives no permission to license the work in any other way, but it does not invalidate such permission if you have separately received it.
d) If the work has interactive user interfaces, each must display Appropriate Legal Notices; however, if the Program has interactive interfaces that do not display Appropriate Legal Notices, your work need not make them do so.
A compilation of a covered work with other separate and independent works, which are not by their nature extensions of the covered work, and which are not combined with it such as to form a larger program, in or on a volume of a storage or distribution medium, is called an "aggregate" if the compilation and its resulting copyright are not used to limit the access or legal rights of the compilation's users beyond what the individual works permit. Inclusion of a covered work in an aggregate does not cause this License to apply to the other parts of the aggregate.
6. Conveying Non-Source Forms.
You may convey a covered work in object code form under the terms of sections 4 and 5, provided that you also convey the machine-readable Corresponding Source under the terms of this License, in one of these ways:
a) Convey the object code in, or embodied in, a physical product (including a physical distribution medium), accompanied by the Corresponding Source fixed on a durable physical medium customarily used for software interchange.
b) Convey the object code in, or embodied in, a physical product (including a physical distribution medium), accompanied by a written offer, valid for at least three years and valid for as long as you offer spare parts or customer support for that product model, to give anyone who possesses the object code either (1) a copy of the Corresponding Source for all the software in the product that is covered by this License, on a durable physical medium customarily used for software interchange, for a price no more than your reasonable cost of physically performing this conveying of source, or (2) access to copy the Corresponding Source from a network server at no charge.
c) Convey individual copies of the object code with a copy of the written offer to provide the Corresponding Source. This alternative is allowed only occasionally and noncommercially, and only if you received the object code with such an offer, in accord with subsection 6b.
d) Convey the object code by offering access from a designated place (gratis or for a charge), and offer equivalent access to the Corresponding Source in the same way through the same place at no further charge. You need not require recipients to copy the Corresponding Source along with the object code. If the place to copy the object code is a network server, the Corresponding Source may be on a different server (operated by you or a third party) that supports equivalent copying facilities, provided you maintain clear directions next to the object code saying where to find the Corresponding Source. Regardless of what server hosts the Corresponding Source, you remain obligated to ensure that it is available for as long as needed to satisfy these requirements.
e) Convey the object code using peer-to-peer transmission, provided you inform other peers where the object code and Corresponding Source of the work are being offered to the general public at no charge under subsection 6d.
A separable portion of the object code, whose source code is excluded from the Corresponding Source as a System Library, need not be included in conveying the object code work.
A "User Product" is either (1) a "consumer product", which means any tangible personal property which is normally used for personal, family, or household purposes, or (2) anything designed or sold for incorporation into a dwelling. In determining whether a product is a consumer product, doubtful cases shall be resolved in favor of coverage. For a particular product received by a particular user, "normally used" refers to a typical or common use of that class of product, regardless of the status of the particular user or of the way in which the particular user actually uses, or expects or is expected to use, the product. A product is a consumer product regardless of whether the product has substantial commercial, industrial or non-consumer uses, unless such uses represent the only significant mode of use of the product.
"Installation Information" for a User Product means any methods, procedures, authorization keys, or other information required to install and execute modified versions of a covered work in that User Product from a modified version of its Corresponding Source. The information must suffice to ensure that the continued functioning of the modified object code is in no case prevented or interfered with solely because modification has been made.
If you convey an object code work under this section in, or with, or specifically for use in, a User Product, and the conveying occurs as part of a transaction in which the right of possession and use of the User Product is transferred to the recipient in perpetuity or for a fixed term (regardless of how the transaction is characterized), the Corresponding Source conveyed under this section must be accompanied by the Installation Information. But this requirement does not apply if neither you nor any third party retains the ability to install modified object code on the User Product (for example, the work has been installed in ROM).
The requirement to provide Installation Information does not include a requirement to continue to provide support service, warranty, or updates for a work that has been modified or installed by the recipient, or for the User Product in which it has been modified or installed. Access to a network may be denied when the modification itself materially and adversely affects the operation of the network or violates the rules and protocols for communication across the network.
Corresponding Source conveyed, and Installation Information provided, in accord with this section must be in a format that is publicly documented (and with an implementation available to the public in source code form), and must require no special password or key for unpacking, reading or copying.
7. Additional Terms.
"Additional permissions" are terms that supplement the terms of this License by making exceptions from one or more of its conditions. Additional permissions that are applicable to the entire Program shall be treated as though they were included in this License, to the extent that they are valid under applicable law. If additional permissions apply only to part of the Program, that part may be used separately under those permissions, but the entire Program remains governed by this License without regard to the additional permissions.
When you convey a copy of a covered work, you may at your option remove any additional permissions from that copy, or from any part of it. (Additional permissions may be written to require their own removal in certain cases when you modify the work.) You may place additional permissions on material, added by you to a covered work, for which you have or can give appropriate copyright permission.
Notwithstanding any other provision of this License, for material you add to a covered work, you may (if authorized by the copyright holders of that material) supplement the terms of this License with terms:
a) Disclaiming warranty or limiting liability differently from the terms of sections 15 and 16 of this License; or
b) Requiring preservation of specified reasonable legal notices or author attributions in that material or in the Appropriate Legal Notices displayed by works containing it; or
c) Prohibiting misrepresentation of the origin of that material, or requiring that modified versions of such material be marked in reasonable ways as different from the original version; or
d) Limiting the use for publicity purposes of names of licensors or authors of the material; or
e) Declining to grant rights under trademark law for use of some trade names, trademarks, or service marks; or
f) Requiring indemnification of licensors and authors of that material by anyone who conveys the material (or modified versions of it) with contractual assumptions of liability to the recipient, for any liability that these contractual assumptions directly impose on those licensors and authors.
All other non-permissive additional terms are considered "further restrictions" within the meaning of section 10. If the Program as you received it, or any part of it, contains a notice stating that it is governed by this License along with a term that is a further restriction, you may remove that term. If a license document contains a further restriction but permits relicensing or conveying under this License, you may add to a covered work material governed by the terms of that license document, provided that the further restriction does not survive such relicensing or conveying.
If you add terms to a covered work in accord with this section, you must place, in the relevant source files, a statement of the additional terms that apply to those files, or a notice indicating where to find the applicable terms.
Additional terms, permissive or non-permissive, may be stated in the form of a separately written license, or stated as exceptions; the above requirements apply either way.
8. Termination.
You may not propagate or modify a covered work except as expressly provided under this License. Any attempt otherwise to propagate or modify it is void, and will automatically terminate your rights under this License (including any patent licenses granted under the third paragraph of section 11).
However, if you cease all violation of this License, then your license from a particular copyright holder is reinstated (a) provisionally, unless and until the copyright holder explicitly and finally terminates your license, and (b) permanently, if the copyright holder fails to notify you of the violation by some reasonable means prior to 60 days after the cessation.
Moreover, your license from a particular copyright holder is reinstated permanently if the copyright holder notifies you of the violation by some reasonable means, this is the first time you have received notice of violation of this License (for any work) from that copyright holder, and you cure the violation prior to 30 days after your receipt of the notice.
Termination of your rights under this section does not terminate the licenses of parties who have received copies or rights from you under this License. If your rights have been terminated and not permanently reinstated, you do not qualify to receive new licenses for the same material under section 10.
9. Acceptance Not Required for Having Copies.
You are not required to accept this License in order to receive or run a copy of the Program. Ancillary propagation of a covered work occurring solely as a consequence of using peer-to-peer transmission to receive a copy likewise does not require acceptance. However, nothing other than this License grants you permission to propagate or modify any covered work. These actions infringe copyright if you do not accept this License. Therefore, by modifying or propagating a covered work, you indicate your acceptance of this License to do so.
10. Automatic Licensing of Downstream Recipients.
Each time you convey a covered work, the recipient automatically receives a license from the original licensors, to run, modify and propagate that work, subject to this License. You are not responsible for enforcing compliance by third parties with this License.
An "entity transaction" is a transaction transferring control of an organization, or substantially all assets of one, or subdividing an organization, or merging organizations. If propagation of a covered work results from an entity transaction, each party to that transaction who receives a copy of the work also receives whatever licenses to the work the party's predecessor in interest had or could give under the previous paragraph, plus a right to possession of the Corresponding Source of the work from the predecessor in interest, if the predecessor has it or can get it with reasonable efforts.
You may not impose any further restrictions on the exercise of the rights granted or affirmed under this License. For example, you may not impose a license fee, royalty, or other charge for exercise of rights granted under this License, and you may not initiate litigation (including a cross-claim or counterclaim in a lawsuit) alleging that any patent claim is infringed by making, using, selling, offering for sale, or importing the Program or any portion of it.
11. Patents.
A "contributor" is a copyright holder who authorizes use under this License of the Program or a work on which the Program is based. The work thus licensed is called the contributor's "contributor version".
A contributor's "essential patent claims" are all patent claims owned or controlled by the contributor, whether already acquired or hereafter acquired, that would be infringed by some manner, permitted by this License, of making, using, or selling its contributor version, but do not include claims that would be infringed only as a consequence of further modification of the contributor version. For purposes of this definition, "control" includes the right to grant patent sublicenses in a manner consistent with the requirements of this License.
Each contributor grants you a non-exclusive, worldwide, royalty-free patent license under the contributor's essential patent claims, to make, use, sell, offer for sale, import and otherwise run, modify and propagate the contents of its contributor version.
In the following three paragraphs, a "patent license" is any express agreement or commitment, however denominated, not to enforce a patent (such as an express permission to practice a patent or covenant not to sue for patent infringement). To "grant" such a patent license to a party means to make such an agreement or commitment not to enforce a patent against the party.
If you convey a covered work, knowingly relying on a patent license, and the Corresponding Source of the work is not available for anyone to copy, free of charge and under the terms of this License, through a publicly available network server or other readily accessible means, then you must either (1) cause the Corresponding Source to be so available, or (2) arrange to deprive yourself of the benefit of the patent license for this particular work, or (3) arrange, in a manner consistent with the requirements of this License, to extend the patent
license to downstream recipients. "Knowingly relying" means you have actual knowledge that, but for the patent license, your conveying the covered work in a country, or your recipient's use of the covered work in a country, would infringe one or more identifiable patents in that country that you have reason to believe are valid.
If, pursuant to or in connection with a single transaction or arrangement, you convey, or propagate by procuring conveyance of, a covered work, and grant a patent license to some of the parties receiving the covered work authorizing them to use, propagate, modify or convey a specific copy of the covered work, then the patent license you grant is automatically extended to all recipients of the covered work and works based on it.
A patent license is "discriminatory" if it does not include within the scope of its coverage, prohibits the exercise of, or is conditioned on the non-exercise of one or more of the rights that are specifically granted under this License. You may not convey a covered work if you are a party to an arrangement with a third party that is in the business of distributing software, under which you make payment to the third party based on the extent of your activity of conveying the work, and under which the third party grants, to any of the parties who would receive the covered work from you, a discriminatory patent license (a) in connection with copies of the covered work conveyed by you (or copies made from those copies), or (b) primarily for and in connection with specific products or compilations that contain the covered work, unless you entered into that arrangement, or that patent license was granted, prior to 28 March 2007.
Nothing in this License shall be construed as excluding or limiting any implied license or other defenses to infringement that may otherwise be available to you under applicable patent law.
12. No Surrender of Others' Freedom.
If conditions are imposed on you (whether by court order, agreement or otherwise) that contradict the conditions of this License, they do not excuse you from the conditions of this License. If you cannot convey a covered work so as to satisfy simultaneously your obligations under this License and any other pertinent obligations, then as a consequence you may
not convey it at all. For example, if you agree to terms that obligate you to collect a royalty for further conveying from those to whom you convey the Program, the only way you could satisfy both those terms and this License would be to refrain entirely from conveying the Program.
13. Remote Network Interaction; Use with the GNU General Public License.
Notwithstanding any other provision of this License, if you modify the Program, your modified version must prominently offer all users interacting with it remotely through a computer network (if your version supports such interaction) an opportunity to receive the Corresponding Source of your version by providing access to the Corresponding Source from a network server at no charge, through some standard or customary means of facilitating copying of software. This Corresponding Source shall include the Corresponding Source for any work covered by version 3 of the GNU General Public License that is incorporated pursuant to the following paragraph.
Notwithstanding any other provision of this License, you have permission to link or combine any covered work with a work licensed under version 3 of the GNU General Public License into a single combined work, and to convey the resulting work. The terms of this License will continue to apply to the part which is the covered work, but the work with which it is combined will remain governed by version 3 of the GNU General Public License.
14. Revised Versions of this License.
The Free Software Foundation may publish revised and/or new versions of the GNU Affero General Public License from time to time. Such new versions will be similar in spirit to the present version, but may differ in detail to address new problems or concerns.
Each version is given a distinguishing version number. If the Program specifies that a certain numbered version of the GNU Affero General Public License "or any later version" applies to it, you have the option of following the terms and conditions either of that numbered version or of any later version published by the Free Software Foundation. If the Program does not specify a version number of the GNU Affero General Public License, you may choose any version ever published by the Free Software Foundation.
If the Program specifies that a proxy can decide which future versions of the GNU Affero General Public License can be used, that proxy's public statement of acceptance of a version permanently authorizes you to choose that version for the Program.
Later license versions may give you additional or different permissions. However, no additional obligations are imposed on any author or copyright holder as a result of your choosing to follow a later version.
15. Disclaimer of Warranty.
THERE IS NO WARRANTY FOR THE PROGRAM, TO THE EXTENT PERMITTED BY APPLICABLE LAW. EXCEPT WHEN OTHERWISE STATED IN WRITING THE COPYRIGHT HOLDERS AND/OR OTHER PARTIES PROVIDE THE PROGRAM "AS IS" WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESSED OR IMPLIED, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE ENTIRE RISK AS TO THE QUALITY AND PERFORMANCE OF THE PROGRAM IS WITH YOU. SHOULD THE PROGRAM PROVE DEFECTIVE, YOU ASSUME THE COST OF ALL NECESSARY SERVICING, REPAIR OR CORRECTION.
16. Limitation of Liability.
IN NO EVENT UNLESS REQUIRED BY APPLICABLE LAW OR AGREED TO IN WRITING WILL ANY COPYRIGHT HOLDER, OR ANY OTHER PARTY WHO MODIFIES AND/OR CONVEYS THE PROGRAM AS PERMITTED ABOVE, BE LIABLE TO YOU FOR DAMAGES, INCLUDING ANY GENERAL, SPECIAL, INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF THE USE OR INABILITY TO USE THE PROGRAM (INCLUDING BUT NOT LIMITED TO LOSS OF DATA OR DATA BEING RENDERED INACCURATE OR LOSSES SUSTAINED BY YOU OR THIRD PARTIES OR A FAILURE OF THE PROGRAM TO OPERATE WITH ANY OTHER PROGRAMS), EVEN IF SUCH HOLDER OR OTHER PARTY HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES.
17. Interpretation of Sections 15 and 16.
If the disclaimer of warranty and limitation of liability provided above cannot be given local legal effect according to their terms, reviewing courts shall apply local law that most closely approximates an absolute waiver of all civil liability in connection with the Program, unless a warranty or assumption of liability accompanies a copy of the Program in return for a fee.
END OF TERMS AND CONDITIONS
How to Apply These Terms to Your New Programs
If you develop a new program, and you want it to be of the greatest possible use to the public, the best way to achieve this is to make it free software which everyone can redistribute and change under these terms.
To do so, attach the following notices to the program. It is safest to attach them to the start of each source file to most effectively state the exclusion of warranty; and each file should have at least the "copyright" line and a pointer to where the full notice is found.
FetchApply
Copyright (C) 2024 VpTechnicalOperations
This program is free software: you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License along with this program. If not, see <http://www.gnu.org/licenses/>.
Also add information on how to contact you by electronic and paper mail.
If your software can interact with users remotely through a computer network, you should also make sure that it provides a way for users to get its source. For example, if your program is a web application, its interface could display a "Source" link that leads users to an archive of the code. There are many ways you could offer source, and different solutions will be better for different programs; see section 13 for the specific requirements.
You should also get your employer (if you work as a programmer) or school, if any, to sign a "copyright disclaimer" for the program, if necessary. For more information on this, and how to apply and follow the GNU AGPL, see <http://www.gnu.org/licenses/>.
@@ -0,0 +1,176 @@
# TSYS FetchApply Testing Framework
## Overview
This testing framework provides comprehensive validation for the TSYS FetchApply infrastructure provisioning system. It includes unit tests, integration tests, security tests, and system validation.
## Test Categories
### 1. Unit Tests (`unit/`)
- **Purpose:** Test individual framework functions and components
- **Scope:** Framework includes, helper functions, syntax validation
- **Example:** `framework-functions.sh` - Tests logging, pretty print, and error handling functions
### 2. Integration Tests (`integration/`)
- **Purpose:** Test complete workflows and module interactions
- **Scope:** End-to-end deployment scenarios, module integration
- **Future:** Module interaction testing, deployment workflow validation
### 3. Security Tests (`security/`)
- **Purpose:** Validate security configurations and practices
- **Scope:** HTTPS enforcement, deployment security, SSH hardening
- **Example:** `https-enforcement.sh` - Validates all URLs use HTTPS
### 4. Validation Tests (`validation/`)
- **Purpose:** System compatibility and pre-flight checks
- **Scope:** System requirements, network connectivity, permissions
- **Example:** `system-requirements.sh` - Validates minimum system requirements
## Usage
### Run All Tests
```bash
./Project-Tests/run-tests.sh
```
### Run Specific Test Categories
```bash
./Project-Tests/run-tests.sh unit # Unit tests only
./Project-Tests/run-tests.sh integration # Integration tests only
./Project-Tests/run-tests.sh security # Security tests only
./Project-Tests/run-tests.sh validation # Validation tests only
```
### Run Individual Tests
```bash
./Project-Tests/validation/system-requirements.sh
./Project-Tests/security/https-enforcement.sh
./Project-Tests/unit/framework-functions.sh
```
## Test Results
- **Console Output:** Real-time test results with color-coded status
- **JSON Reports:** Detailed test reports saved to `logs/tests/`
- **Exit Codes:** 0 for success, 1 for failures
## Configuration Validation
The validation framework performs pre-flight checks to ensure system compatibility:
### System Requirements
- **Memory:** Minimum 2GB RAM
- **Disk Space:** Minimum 10GB available
- **OS Compatibility:** Ubuntu/Debian (tested), others (may work)
### Network Connectivity
- Tests connection to required download sources
- Validates HTTPS endpoints are accessible
- Checks for firewall/proxy issues
### Command Dependencies
- Verifies required tools are installed (`curl`, `wget`, `git`, `systemctl`, `apt-get`)
- Checks for proper versions where applicable
### Permissions
- Validates write access to system directories
- Checks for required administrative privileges
## Adding New Tests
### Test File Structure
```bash
#!/bin/bash
set -euo pipefail
function test_something() {
echo "🔍 Testing something..."
if [[ condition ]]; then
echo "✅ Test passed"
return 0
else
echo "❌ Test failed"
return 1
fi
}
function main() {
echo "🧪 Running Test Suite Name"
echo "=========================="
local total_failures=0
test_something || ((total_failures++))
echo "=========================="
if [[ $total_failures -eq 0 ]]; then
echo "✅ All tests passed"
exit 0
else
echo "$total_failures tests failed"
exit 1
fi
}
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
```
### Test Categories Guidelines
- **Unit Tests:** Focus on individual functions, fast execution
- **Integration Tests:** Test module interactions, longer execution
- **Security Tests:** Validate security configurations
- **Validation Tests:** Pre-flight system checks
## Continuous Integration
The testing framework is designed to integrate with CI/CD pipelines:
```bash
# Example CI script
./Project-Tests/run-tests.sh all
test_exit_code=$?
if [[ $test_exit_code -eq 0 ]]; then
echo "All tests passed - deployment approved"
else
echo "Tests failed - deployment blocked"
exit 1
fi
```
## Test Development Best Practices
1. **Clear Test Names:** Use descriptive function names
2. **Proper Exit Codes:** Return 0 for success, 1 for failure
3. **Informative Output:** Use emoji and clear messages
4. **Timeout Protection:** Use timeout for network operations
5. **Cleanup:** Remove temporary files and resources
6. **Error Handling:** Use `set -euo pipefail` for strict error handling
## Troubleshooting
### Common Issues
- **Permission Denied:** Run tests with appropriate privileges
- **Network Timeouts:** Check firewall and proxy settings
- **Missing Dependencies:** Install required tools before testing
- **Script Errors:** Validate syntax with `bash -n script.sh`
### Debug Mode
```bash
# Enable debug output
export DEBUG=1
./Project-Tests/run-tests.sh
```
## Contributing
When adding new functionality to FetchApply:
1. Add corresponding tests in appropriate category
2. Run full test suite before committing
3. Update documentation for new test cases
4. Ensure tests pass in clean environment
+97
View File
@@ -0,0 +1,97 @@
#!/usr/bin/bash
#
# remote.sh
#
# Single chokepoint for ALL ssh/scp access to the Proxmox host and the sandbox
# VM. Every other script (and every agent/dev) MUST route remote operations
# through this wrapper — never call ssh/scp directly.
#
# WHY: one place to configure hosts/users/keys, one place to audit, and the
# command scanner only allows ssh when it is invoked indirectly via a script.
#
# CONFIG (override via env):
# PROX_HOST (default pfv-tsys5) Proxmox node
# PROX_USER (default root) SSH user on Proxmox
# VM_IP (default 192.168.3.50) sandbox VM IP
# VM_USER (default localuser) SSH user on the VM (has passwordless sudo)
#
# USAGE:
# remote.sh prox <cmd...> run command on Proxmox
# remote.sh vm <cmd...> run command on VM as $VM_USER
# remote.sh vmroot <cmd...> run command on VM as root via sudo
# remote.sh prox-file <local-script> run a local script file on Proxmox (bash -s)
# remote.sh vm-file <local-script> run a local script file on the VM (bash -s)
# remote.sh vm-copy <local> <dest> copy a local file to the VM (~$VM_USER space)
# remote.sh prox-copy <local> <dest> copy a local file to Proxmox
#
set -uo pipefail
PROX_HOST="${PROX_HOST:-pfv-tsys5}"
PROX_USER="${PROX_USER:-root}"
VM_IP="${VM_IP:-192.168.3.50}"
VM_USER="${VM_USER:-localuser}"
VM_ID="${VM_ID:-}"
GUEST_TIMEOUT="${GUEST_TIMEOUT:-900}"
SSH_OPTS=(-o BatchMode=yes -o StrictHostKeyChecking=accept-new -o ConnectTimeout=15)
die() { echo "remote.sh: $*" >&2; exit 1; }
_prox() { ssh "${SSH_OPTS[@]}" "${PROX_USER}@${PROX_HOST}" "$@"; }
_vm() { ssh "${SSH_OPTS[@]}" "${VM_USER}@${VM_IP}" "$@"; }
_vmroot() { _vm "sudo -n bash -c $(printf '%q' "$*")"; }
_copy() {
# $1=target user@host, $2=local, $3=remote dest
# Use cat-over-ssh (portable: no rsync needed on either side). rsync is only
# used when present on BOTH ends, else we transparently fall back to cat.
local target="$1" local="$2" dest="$3"
local userhost="${target%@*}@${target#*@}"
if command -v rsync >/dev/null 2>&1 \
&& ssh "${SSH_OPTS[@]}" "$userhost" 'command -v rsync' >/dev/null 2>&1; then
rsync -az -e "ssh ${SSH_OPTS[*]}" "$local" "${userhost}:${dest}"
else
ssh "${SSH_OPTS[@]}" "$userhost" "cat > '$dest'" < "$local"
fi
}
# Out-of-band VM access via the Proxmox qemu-guest-agent. This runs commands
# as root inside the VM and does NOT depend on SSH, so it works even after
# secharden-ssh replaces authorized_keys and secharden-2fa enforces
# publickey+keyboard-interactive (which blocks non-interactive SSH).
GUEST_PARSER="/root/.knel-guest-parse.py"
GUEST_PARSER_SRC="import sys, json
try:
d = json.load(sys.stdin)
except Exception:
sys.exit(3)
sys.stdout.write(d.get('out-data', '') or '')
sys.stderr.write(d.get('err-data', '') or '')
ec = d.get('exitcode', 1)
sys.exit(ec if ec is not None else 1)"
_ensure_guest_parser() {
if _prox "test -f '$GUEST_PARSER'" >/dev/null 2>&1; then return 0; fi
printf '%s\n' "$GUEST_PARSER_SRC" | _prox "cat > '$GUEST_PARSER'" >/dev/null 2>&1
}
_vm_guest() {
[ -n "$VM_ID" ] || die "vm-guest requires VM_ID"
_ensure_guest_parser
local cmdb64; cmdb64="$(printf '%s' "$*" | base64 -w0)"
_prox "qm guest exec $VM_ID --timeout ${GUEST_TIMEOUT} -- /bin/sh -c 'echo $cmdb64 | base64 -d | /bin/sh' 2>/dev/null | python3 '$GUEST_PARSER'"
}
mode="${1:-}"; shift || true
case "$mode" in
prox) [ "$#" -ge 0 ] || die "need command"; _prox "$*" ;;
vm) _vm "$*" ;;
vmroot) [ "$#" -ge 1 ] || die "need command"; _vmroot "$*" ;;
prox-file) [ -f "${1:-}" ] || die "need local script file"; _prox "bash -s" < "$1" ;;
vm-file) [ -f "${1:-}" ] || die "need local script file"; _vm "bash -s" < "$1" ;;
vm-copy) [ -f "${1:-}" ] || die "need local file"; _copy "${VM_USER}@${VM_IP}" "$1" "${2:-}" ;;
prox-copy) [ -f "${1:-}" ] || die "need local file"; _copy "${PROX_USER}@${PROX_HOST}" "$1" "${2:-}" ;;
vm-guest) [ "$#" -ge 1 ] || die "need command"; _vm_guest "$*" ;;
""|-h|--help|help) sed -n '2,40p' "${BASH_SOURCE[0]}" >&2; exit 0 ;;
*) die "unknown mode '$mode'. Run '$0 help'." ;;
esac
+137
View File
@@ -0,0 +1,137 @@
#!/bin/bash
# TSYS FetchApply Testing Framework
# Main test runner script
set -euo pipefail
# Resolve repository root from this script's location (Project-Tests/ -> repo root)
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
# The KNELShellFramework is vendored under vendor/
FRAMEWORK_INCLUDES="$PROJECT_ROOT/vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework/Framework-Includes"
source "$FRAMEWORK_INCLUDES/Logging.sh"
source "$FRAMEWORK_INCLUDES/PrettyPrint.sh"
# The vendored PrettyPrint only defines print_info/print_error; provide the
# additional output helpers the test suite relies on.
function print_header() { echo ""; echo "=== $1 ==="; }
function print_success() { echo "$1"; }
function print_warning() { echo "⚠️ $1"; }
# Test configuration
TEST_LOG_DIR="$PROJECT_ROOT/logs/tests"
TEST_RESULTS_FILE="$TEST_LOG_DIR/test-results-$(date +%Y%m%d-%H%M%S).json"
# Ensure test log directory exists
mkdir -p "$TEST_LOG_DIR"
# Test counters
declare -g TESTS_PASSED=0
declare -g TESTS_FAILED=0
declare -g TESTS_SKIPPED=0
# Test runner functions
function run_test_suite() {
local suite_name="$1"
local test_dir="$2"
print_header "Running $suite_name Tests"
if [[ ! -d "$test_dir" ]]; then
print_warning "Test directory $test_dir not found, skipping"
return 0
fi
for test_file in "$test_dir"/*.sh; do
if [[ -f "$test_file" ]]; then
run_single_test "$test_file"
fi
done
}
function run_single_test() {
local test_file="$1"
local test_name="$(basename "$test_file" .sh)"
print_info "Running test: $test_name"
if timeout 300 bash "$test_file"; then
print_success "$test_name PASSED"
TESTS_PASSED=$((TESTS_PASSED + 1))
else
print_error "$test_name FAILED"
TESTS_FAILED=$((TESTS_FAILED + 1))
fi
}
function generate_test_report() {
local total_tests=$((TESTS_PASSED + TESTS_FAILED + TESTS_SKIPPED))
print_header "Test Results Summary"
print_info "Total Tests: $total_tests"
print_success "Passed: $TESTS_PASSED"
print_error "Failed: $TESTS_FAILED"
print_warning "Skipped: $TESTS_SKIPPED"
# Generate JSON report
cat > "$TEST_RESULTS_FILE" <<EOF
{
"timestamp": "$(date -Iseconds)",
"total_tests": $total_tests,
"passed": $TESTS_PASSED,
"failed": $TESTS_FAILED,
"skipped": $TESTS_SKIPPED,
"success_rate": $(awk "BEGIN {printf \"%.2f\", ($TESTS_PASSED/$total_tests)*100}")
}
EOF
print_info "Test report saved to: $TEST_RESULTS_FILE"
}
# Main execution
function main() {
print_header "TSYS FetchApply Test Suite"
# Parse command line arguments
local test_type="${1:-all}"
case "$test_type" in
"unit")
run_test_suite "Unit" "$(dirname "$0")/unit"
;;
"integration")
run_test_suite "Integration" "$(dirname "$0")/integration"
;;
"security")
run_test_suite "Security" "$(dirname "$0")/security"
;;
"validation")
run_test_suite "Validation" "$(dirname "$0")/validation"
;;
"all")
run_test_suite "Unit" "$(dirname "$0")/unit"
run_test_suite "Integration" "$(dirname "$0")/integration"
run_test_suite "Security" "$(dirname "$0")/security"
run_test_suite "Validation" "$(dirname "$0")/validation"
;;
*)
print_error "Usage: $0 [unit|integration|security|validation|all]"
exit 1
;;
esac
generate_test_report
# Exit with appropriate code
if [[ $TESTS_FAILED -gt 0 ]]; then
exit 1
else
exit 0
fi
}
# Run main if executed directly
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
@@ -0,0 +1,312 @@
#!/bin/bash
# Two-Factor Authentication Validation Test
# Validates 2FA configuration for SSH, Cockpit, and Webmin
set -euo pipefail
PROJECT_ROOT="$(dirname "$(realpath "${BASH_SOURCE[0]}")")/../.."
function test_2fa_packages() {
echo "🔍 Testing 2FA package installation..."
local packages=("libpam-google-authenticator" "qrencode")
local failed=0
for package in "${packages[@]}"; do
if dpkg -s "$package" 2>/dev/null | grep -q "^Status:.*installed"; then
echo "✅ Package installed: $package"
else
echo "❌ Package missing: $package"
((++failed))
fi
done
# Check if google-authenticator command exists
if command -v google-authenticator >/dev/null 2>&1; then
echo "✅ Google Authenticator command available"
else
echo "❌ Google Authenticator command not found"
((++failed))
fi
return $failed
}
function test_ssh_2fa_config() {
echo "🔍 Testing SSH 2FA configuration..."
local ssh_config="/etc/ssh/sshd_config"
local failed=0
# Check required SSH settings
if grep -q "^ChallengeResponseAuthentication yes" "$ssh_config"; then
echo "✅ ChallengeResponseAuthentication enabled"
else
echo "❌ ChallengeResponseAuthentication not enabled"
((++failed))
fi
if grep -q "^UsePAM yes" "$ssh_config"; then
echo "✅ UsePAM enabled"
else
echo "❌ UsePAM not enabled"
((++failed))
fi
if grep -q "^AuthenticationMethods publickey,keyboard-interactive" "$ssh_config"; then
echo "✅ AuthenticationMethods configured for 2FA"
else
echo "❌ AuthenticationMethods not configured for 2FA"
((++failed))
fi
return $failed
}
function test_pam_2fa_config() {
echo "🔍 Testing PAM 2FA configuration..."
local pam_sshd="/etc/pam.d/sshd"
local failed=0
# Check if PAM includes Google Authenticator
if grep -q "pam_google_authenticator.so" "$pam_sshd"; then
echo "✅ PAM Google Authenticator module configured"
else
echo "❌ PAM Google Authenticator module not configured"
((++failed))
fi
# Check if nullok is present (allows users without 2FA setup)
if grep -q "pam_google_authenticator.so nullok" "$pam_sshd"; then
echo "✅ PAM nullok option configured (allows gradual rollout)"
else
echo "⚠️ PAM nullok option not configured (immediate enforcement)"
fi
return $failed
}
function test_cockpit_2fa_config() {
echo "🔍 Testing Cockpit 2FA configuration..."
local cockpit_config="/etc/cockpit/cockpit.conf"
local cockpit_pam="/etc/pam.d/cockpit"
local failed=0
# Check if Cockpit is installed
if ! command -v cockpit-ws >/dev/null 2>&1; then
echo "⚠️ Cockpit not installed, skipping test"
return 0
fi
# Check Cockpit configuration
if [[ -f "$cockpit_config" ]]; then
echo "✅ Cockpit configuration file exists"
else
echo "❌ Cockpit configuration file missing"
((++failed))
fi
# Check Cockpit PAM configuration
if [[ -f "$cockpit_pam" ]] && grep -q "pam_google_authenticator.so" "$cockpit_pam"; then
echo "✅ Cockpit PAM 2FA configured"
else
echo "❌ Cockpit PAM 2FA not configured"
((++failed))
fi
return $failed
}
function test_webmin_2fa_config() {
echo "🔍 Testing Webmin 2FA configuration..."
local webmin_config="/etc/webmin/miniserv.conf"
local failed=0
# Check if Webmin is installed
if [[ ! -f "$webmin_config" ]]; then
echo "⚠️ Webmin not installed, skipping test"
return 0
fi
# Check Webmin 2FA settings
if grep -q "^twofactor_provider=totp" "$webmin_config"; then
echo "✅ Webmin TOTP provider configured"
else
echo "❌ Webmin TOTP provider not configured"
((++failed))
fi
if grep -q "^twofactor=1" "$webmin_config"; then
echo "✅ Webmin 2FA enabled"
else
echo "❌ Webmin 2FA not enabled"
((++failed))
fi
return $failed
}
function test_user_2fa_setup() {
echo "🔍 Testing user 2FA setup preparation..."
local users=("localuser" "root")
local failed=0
for user in "${users[@]}"; do
if id "$user" &>/dev/null; then
local user_home; user_home="$(getent passwd "$user" | cut -d: -f6)"
# Check if setup script exists
if [[ -f "/tmp/setup-2fa-$user.sh" ]]; then
echo "✅ 2FA setup script exists for user: $user"
else
echo "❌ 2FA setup script missing for user: $user"
((++failed))
fi
# Check if instructions exist
if [[ -n "$user_home" && -f "$user_home/2fa-setup-instructions.txt" ]]; then
echo "✅ 2FA instructions exist for user: $user"
else
echo "❌ 2FA instructions missing for user: $user"
((++failed))
fi
else
echo "⚠️ User $user not found, skipping"
fi
done
return $failed
}
function test_service_status() {
echo "🔍 Testing service status..."
local failed=0
# Test SSH service
if systemctl is-active sshd >/dev/null 2>&1; then
echo "✅ SSH service is running"
else
echo "❌ SSH service is not running"
((++failed))
fi
# Test SSH configuration
if sshd -t 2>/dev/null; then
echo "✅ SSH configuration is valid"
else
echo "❌ SSH configuration is invalid"
((++failed))
fi
# Test Cockpit service if installed
if systemctl is-enabled cockpit.socket >/dev/null 2>&1; then
if systemctl is-active cockpit.socket >/dev/null 2>&1; then
echo "✅ Cockpit service is running"
else
echo "❌ Cockpit service is not running"
((++failed))
fi
fi
# Test Webmin service if installed
if systemctl is-enabled webmin >/dev/null 2>&1; then
if systemctl is-active webmin >/dev/null 2>&1; then
echo "✅ Webmin service is running"
else
echo "❌ Webmin service is not running"
((++failed))
fi
fi
return $failed
}
function test_backup_existence() {
echo "🔍 Testing backup existence..."
local backup_dir="/root/backup"
local failed=0
if [[ -d "$backup_dir" ]]; then
# Look for recent 2FA backups
local recent_backups=$(find "$backup_dir" -name "2fa-*" -type d -newer /etc/ssh/sshd_config 2>/dev/null | wc -l)
if [[ $recent_backups -gt 0 ]]; then
echo "✅ Recent 2FA backup found in $backup_dir"
else
echo "⚠️ No recent 2FA backups found"
fi
else
echo "❌ Backup directory does not exist"
((++failed))
fi
return $failed
}
function test_2fa_enforcement() {
echo "🔍 Testing 2FA enforcement level..."
local pam_sshd="/etc/pam.d/sshd"
# Check if nullok is used (allows users without 2FA)
if grep -q "pam_google_authenticator.so nullok" "$pam_sshd"; then
echo "⚠️ 2FA enforcement: GRADUAL (nullok allows users without 2FA)"
echo " Users can log in without 2FA during setup phase"
else
echo "✅ 2FA enforcement: STRICT (all users must have 2FA)"
echo " All users must have 2FA configured to log in"
fi
return 0
}
# Main test execution
function main() {
echo "🔒 Running Two-Factor Authentication Validation Tests"
echo "=================================================="
local total_failures=0
# Run all 2FA validation tests
test_2fa_packages || ((total_failures++))
test_ssh_2fa_config || ((total_failures++))
test_pam_2fa_config || ((total_failures++))
test_cockpit_2fa_config || ((total_failures++))
test_webmin_2fa_config || ((total_failures++))
test_user_2fa_setup || ((total_failures++))
test_service_status || ((total_failures++))
test_backup_existence || ((total_failures++))
test_2fa_enforcement || ((total_failures++))
echo "=================================================="
if [[ $total_failures -eq 0 ]]; then
echo "✅ All 2FA validation tests passed"
echo ""
echo "📋 Next Steps:"
echo "1. Run user setup scripts: /tmp/setup-2fa-*.sh"
echo "2. Test 2FA login from another terminal"
echo "3. Remove nullok from PAM config for strict enforcement"
exit 0
else
echo "$total_failures 2FA validation tests failed"
echo ""
echo "🔧 Troubleshooting:"
echo "1. Re-run secharden-2fa.sh script"
echo "2. Check system logs: journalctl -u sshd"
echo "3. Verify package installation"
exit 1
fi
}
# Run main if executed directly
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
@@ -0,0 +1,148 @@
#!/bin/bash
# HTTPS Enforcement Security Test
# Validates that all scripts use HTTPS instead of HTTP
set -euo pipefail
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
FRAMEWORK_INCLUDES="$PROJECT_ROOT/vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework/Framework-Includes"
function test_no_http_urls() {
echo "🔍 Checking for HTTP URLs in scripts..."
local http_violations=0
local script_dirs=("$PROJECT_ROOT/ProjectCode" "$FRAMEWORK_INCLUDES" "$PROJECT_ROOT/Project-Includes")
for dir in "${script_dirs[@]}"; do
if [[ -d "$dir" ]]; then
# Find HTTP URLs in shell scripts (excluding comments)
while IFS= read -r -d '' file; do
# grep -n prefixes "linenum:", so the comment filter must allow
# for that prefix before the leading '#' of a comment line.
if grep -n "http://" "$file" | grep -vE '^[0-9]+:[[:space:]]*#' | grep -v "schema.org" | grep -v "xmlns"; then
echo "❌ HTTP URL found in: $file"
((++http_violations))
fi
done < <(find "$dir" -name "*.sh" -type f -print0)
fi
done
if [[ $http_violations -eq 0 ]]; then
echo "✅ No HTTP URLs found in active scripts"
return 0
else
echo "❌ Found $http_violations HTTP URL violations"
return 1
fi
}
function test_https_urls_valid() {
echo "🔍 Validating HTTPS URLs are accessible..."
local script_dirs=("$PROJECT_ROOT/ProjectCode" "$FRAMEWORK_INCLUDES" "$PROJECT_ROOT/Project-Includes")
local https_failures=0
# Extract HTTPS URLs from scripts
for dir in "${script_dirs[@]}"; do
if [[ -d "$dir" ]]; then
while IFS= read -r -d '' file; do
# Extract HTTPS URLs from non-comment lines
grep -o "https://[^[:space:]\"']*" "$file" | grep -v "schema.org" | while read -r url; do
# Test connectivity with timeout
if timeout 30 curl -s --head --fail "$url" >/dev/null 2>&1; then
echo "✅ HTTPS URL accessible: $url"
else
echo "❌ HTTPS URL not accessible: $url"
((++https_failures))
fi
done
done < <(find "$dir" -name "*.sh" -type f -print0)
fi
done
return $https_failures
}
function test_ssl_certificate_validation() {
echo "🔍 Testing SSL certificate validation..."
local test_urls=(
"https://archive.ubuntu.com"
"https://linux.dell.com"
"https://download.proxmox.com"
)
local ssl_failures=0
for url in "${test_urls[@]}"; do
# Verify TLS is required and the certificate chain is valid. Do NOT use
# --cert-status: that requires OCSP stapling, which many valid CDNs do
# not provide, producing false negatives for otherwise-valid certs.
if curl -s --fail --ssl-reqd "$url" >/dev/null 2>&1; then
echo "✅ SSL certificate valid: $url"
else
echo "❌ SSL certificate validation failed: $url"
((++ssl_failures))
fi
done
return $ssl_failures
}
function test_deployment_security() {
echo "🔍 Testing deployment method security..."
local readme_file="$PROJECT_ROOT/README.md"
if [[ -f "$readme_file" ]]; then
# Check for insecure curl | bash patterns
if grep -q "curl.*|.*bash" "$readme_file" || grep -q "wget.*|.*bash" "$readme_file"; then
echo "❌ Insecure deployment method found in README.md"
return 1
else
echo "✅ Secure deployment method in README.md"
fi
# Check for git clone method
if grep -q "git clone" "$readme_file"; then
echo "✅ Git clone deployment method found"
return 0
else
echo "⚠️ No git clone method found in README.md"
return 1
fi
else
echo "❌ README.md not found"
return 1
fi
}
# Main test execution
function main() {
echo "🔒 Running HTTPS Enforcement Security Tests"
echo "=========================================="
local total_failures=0
# Run all security tests
test_no_http_urls || ((total_failures++))
test_https_urls_valid || ((total_failures++))
test_ssl_certificate_validation || ((total_failures++))
test_deployment_security || ((total_failures++))
echo "=========================================="
if [[ $total_failures -eq 0 ]]; then
echo "✅ All HTTPS enforcement security tests passed"
exit 0
else
echo "$total_failures HTTPS enforcement security tests failed"
exit 1
fi
}
# Run main if executed directly
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
@@ -0,0 +1,179 @@
#!/bin/bash
# Framework Functions Unit Tests
# Tests core framework functionality
set -euo pipefail
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
# Source framework functions from the vendored KNELShellFramework
FRAMEWORK_INCLUDES="$PROJECT_ROOT/vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework/Framework-Includes"
source "$FRAMEWORK_INCLUDES/Logging.sh" 2>/dev/null || echo "Warning: Logging.sh not found"
source "$FRAMEWORK_INCLUDES/PrettyPrint.sh" 2>/dev/null || echo "Warning: PrettyPrint.sh not found"
source "$FRAMEWORK_INCLUDES/ErrorHandling.sh" 2>/dev/null || echo "Warning: ErrorHandling.sh not found"
function test_logging_variables() {
echo "🔍 Testing logging variables..."
if [[ -n "${CURRENT_TIMESTAMP:-}" ]]; then
echo "✅ CURRENT_TIMESTAMP is set"
else
echo "❌ CURRENT_TIMESTAMP is not set"
return 1
fi
if [[ -n "${LOGFILENAME:-}" ]]; then
echo "✅ LOGFILENAME is set"
else
echo "❌ LOGFILENAME is not set"
return 1
fi
return 0
}
function test_pretty_print_functions() {
echo "🔍 Testing pretty print functions..."
# Test if pretty print functions exist
if command -v print_info >/dev/null 2>&1; then
print_info "Test info message" >/dev/null 2>&1 || true
echo "✅ print_info function exists"
else
echo "❌ print_info function missing"
return 1
fi
if command -v print_error >/dev/null 2>&1; then
print_error "Test error message" >/dev/null 2>&1 || true
echo "✅ print_error function exists"
else
echo "❌ print_error function missing"
return 1
fi
return 0
}
function test_error_handling() {
echo "🔍 Testing error handling..."
# Test if error handling functions exist
if command -v error_out >/dev/null 2>&1; then
echo "✅ error_out function exists"
else
echo "❌ error_out function missing"
return 1
fi
if command -v handle_failure >/dev/null 2>&1; then
echo "✅ handle_failure function exists"
else
echo "❌ handle_failure function missing"
return 1
fi
# Test bash strict mode is set
if [[ "$-" == *e* ]]; then
echo "✅ Bash strict mode (set -e) is enabled"
else
echo "❌ Bash strict mode (set -e) not enabled"
return 1
fi
if [[ "$-" == *u* ]]; then
echo "✅ Bash unset variable checking (set -u) is enabled"
else
echo "❌ Bash unset variable checking (set -u) not enabled"
return 1
fi
return 0
}
function test_framework_includes_exist() {
echo "🔍 Testing framework includes exist..."
local required_includes=(
"Logging.sh"
"PrettyPrint.sh"
"ErrorHandling.sh"
"PreflightCheck.sh"
)
local missing_files=0
for include_file in "${required_includes[@]}"; do
if [[ -f "$FRAMEWORK_INCLUDES/$include_file" ]]; then
echo "✅ Framework include exists: $include_file"
else
echo "❌ Framework include missing: $include_file"
((++missing_files))
fi
done
return $missing_files
}
function test_syntax_validation() {
echo "🔍 Testing script syntax validation..."
local syntax_errors=0
local script_dirs=(
"$FRAMEWORK_INCLUDES"
"$PROJECT_ROOT/Project-Includes"
"$PROJECT_ROOT/ProjectCode"
)
for dir in "${script_dirs[@]}"; do
if [[ -d "$dir" ]]; then
while IFS= read -r -d '' file; do
# Skip files that aren't bash scripts despite a .sh extension (e.g. PHP agents)
local shebang
shebang="$(head -c 32 "$file" 2>/dev/null)"
case "$shebang" in
*php*|*python*|*perl*) continue ;;
esac
if bash -n "$file" 2>/dev/null; then
echo "✅ Syntax valid: $(basename "$file")"
else
echo "❌ Syntax error in: $(basename "$file")"
((++syntax_errors))
fi
done < <(find "$dir" -name "*.sh" -type f -print0)
fi
done
return $syntax_errors
}
# Main test execution
function main() {
echo "🧪 Running Framework Functions Unit Tests"
echo "========================================"
local total_failures=0
# Run all unit tests
test_framework_includes_exist || ((total_failures++))
test_logging_variables || ((total_failures++))
test_pretty_print_functions || ((total_failures++))
test_error_handling || ((total_failures++))
test_syntax_validation || ((total_failures++))
echo "========================================"
if [[ $total_failures -eq 0 ]]; then
echo "✅ All framework function unit tests passed"
exit 0
else
echo "$total_failures framework function unit tests failed"
exit 1
fi
}
# Run main if executed directly
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
+297
View File
@@ -0,0 +1,297 @@
#!/bin/bash
# Safe Download Framework Unit Tests
# Tests the SafeDownload.sh framework functionality
set -euo pipefail
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
# Source framework functions from the vendored KNELShellFramework
FRAMEWORK_INCLUDES="$PROJECT_ROOT/vendor/git@git.knownelement.com/29418/KNEL/KNELShellFramework/Framework-Includes"
# The vendored PrettyPrint only defines print_info/print_error, but SafeDownload.sh
# calls print_success/print_warning; define lightweight shims before sourcing.
function print_success() { echo "$1"; }
function print_warning() { echo "⚠️ $1"; }
source "$FRAMEWORK_INCLUDES/SafeDownload.sh"
function test_network_connectivity() {
echo "🔍 Testing network connectivity..."
if check_url_accessibility "https://github.com"; then
echo "✅ Network connectivity test passed"
return 0
else
echo "❌ Network connectivity test failed"
return 1
fi
}
function test_url_accessibility() {
echo "🔍 Testing URL accessibility..."
local test_urls=(
"https://archive.ubuntu.com"
"https://github.com"
)
local failed=0
for url in "${test_urls[@]}"; do
if check_url_accessibility "$url"; then
echo "✅ URL accessible: $url"
else
echo "❌ URL not accessible: $url"
((++failed))
fi
done
return $failed
}
function test_safe_download() {
echo "🔍 Testing safe download functionality..."
local test_url="https://raw.githubusercontent.com/torvalds/linux/master/README"
local test_dest="/tmp/test-download-$$"
local failed=0
# Test successful download
if safe_download "$test_url" "$test_dest"; then
echo "✅ Safe download successful"
# Verify file exists and has content
if [[ -f "$test_dest" && -s "$test_dest" ]]; then
echo "✅ Downloaded file exists and has content"
else
echo "❌ Downloaded file is missing or empty"
((++failed))
fi
# Cleanup
rm -f "$test_dest"
else
echo "❌ Safe download failed"
((++failed))
fi
# Test download with invalid URL
if safe_download "https://invalid.example.com/nonexistent" "/tmp/test-invalid-$$" 2>/dev/null; then
echo "❌ Invalid URL download should have failed"
((++failed))
else
echo "✅ Invalid URL download failed as expected"
fi
return $failed
}
function test_checksum_verification() {
echo "🔍 Testing checksum verification..."
local test_file="/tmp/test-checksum-$$"
local test_content="Hello, World!"
local expected_checksum="dffd6021bb2bd5b0af676290809ec3a53191dd81c7f70a4b28688a362182986f"
local failed=0
# Create test file with known content
echo -n "$test_content" > "$test_file"
# Test correct checksum
if verify_checksum "$test_file" "$expected_checksum"; then
echo "✅ Correct checksum verification passed"
else
echo "❌ Correct checksum verification failed"
((++failed))
fi
# Test incorrect checksum
if verify_checksum "$test_file" "invalid_checksum" 2>/dev/null; then
echo "❌ Incorrect checksum should have failed"
((++failed))
else
echo "✅ Incorrect checksum verification failed as expected"
fi
# Test missing file
if verify_checksum "/tmp/nonexistent-file-$$" "$expected_checksum" 2>/dev/null; then
echo "❌ Missing file checksum should have failed"
((++failed))
else
echo "✅ Missing file checksum verification failed as expected"
fi
# Cleanup
rm -f "$test_file"
return $failed
}
function test_batch_download() {
echo "🔍 Testing batch download functionality..."
# Create test download map
declare -A test_downloads=(
["https://raw.githubusercontent.com/torvalds/linux/master/README"]="/tmp/batch-test-1-$$"
["https://raw.githubusercontent.com/torvalds/linux/master/COPYING"]="/tmp/batch-test-2-$$"
)
local failed=0
# Test batch download
if batch_download test_downloads; then
echo "✅ Batch download successful"
# Verify all files were downloaded
for file in "${test_downloads[@]}"; do
if [[ -f "$file" && -s "$file" ]]; then
echo "✅ Batch file downloaded: $(basename "$file")"
else
echo "❌ Batch file missing: $(basename "$file")"
((++failed))
fi
done
# Cleanup
for file in "${test_downloads[@]}"; do
rm -f "$file"
done
else
echo "❌ Batch download failed"
((++failed))
fi
return $failed
}
function test_config_backup_and_restore() {
echo "🔍 Testing config backup and restore..."
local test_config="/tmp/test-config-$$"
local original_content="Original configuration"
local failed=0
# Create original config file
echo "$original_content" > "$test_config"
# Test safe config download (this will fail with invalid URL, triggering restore)
if safe_config_download "https://invalid.example.com/config" "$test_config" ".test-backup" 2>/dev/null; then
echo "❌ Invalid config download should have failed"
((++failed))
else
echo "✅ Invalid config download failed as expected"
# Verify original file was restored
if [[ -f "$test_config" ]] && grep -q "$original_content" "$test_config"; then
echo "✅ Original config was restored after failed download"
else
echo "❌ Original config was not restored properly"
((++failed))
fi
fi
# Cleanup
rm -f "$test_config" "$test_config.test-backup"
return $failed
}
function test_download_error_handling() {
echo "🔍 Testing download error handling..."
local failed=0
# Test download with missing parameters
if safe_download "" "/tmp/test" 2>/dev/null; then
echo "❌ Download with empty URL should have failed"
((++failed))
else
echo "✅ Download with empty URL failed as expected"
fi
if safe_download "https://example.com" "" 2>/dev/null; then
echo "❌ Download with empty destination should have failed"
((++failed))
else
echo "✅ Download with empty destination failed as expected"
fi
# Test download to read-only location (should fail). Only meaningful for
# non-root users: root bypasses filesystem permissions, so the expected
# write failure never happens and the assertion is invalid.
if [[ $EUID -eq 0 ]]; then
echo "⏭️ Skipping read-only-location test (running as root; root bypasses FS perms)"
elif safe_download "https://github.com" "/test-readonly-$$" 2>/dev/null; then
echo "❌ Download to read-only location should have failed"
((++failed))
else
echo "✅ Download to read-only location failed as expected"
fi
return $failed
}
function test_download_performance() {
echo "🔍 Testing download performance..."
local test_url="https://raw.githubusercontent.com/torvalds/linux/master/README"
local test_dest="/tmp/perf-test-$$"
local start_time end_time duration
start_time=$(date +%s)
if safe_download "$test_url" "$test_dest"; then
end_time=$(date +%s)
duration=$((end_time - start_time))
echo "✅ Download completed in ${duration}s"
if [[ $duration -gt 30 ]]; then
echo "⚠️ Download took longer than expected (>30s)"
else
echo "✅ Download performance acceptable"
fi
# Cleanup
rm -f "$test_dest"
return 0
else
echo "❌ Performance test download failed"
return 1
fi
}
# Main test execution
function main() {
echo "🧪 Running Safe Download Framework Unit Tests"
echo "==========================================="
local total_failures=0
# Run all tests
test_network_connectivity || ((total_failures++))
test_url_accessibility || ((total_failures++))
test_safe_download || ((total_failures++))
test_checksum_verification || ((total_failures++))
test_batch_download || ((total_failures++))
test_config_backup_and_restore || ((total_failures++))
test_download_error_handling || ((total_failures++))
test_download_performance || ((total_failures++))
echo "==========================================="
if [[ $total_failures -eq 0 ]]; then
echo "✅ All safe download framework tests passed"
exit 0
else
echo "$total_failures safe download framework tests failed"
exit 1
fi
}
# Run main if executed directly
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
@@ -0,0 +1,202 @@
#!/bin/bash
# Redundant DNS/NTP Validation Test
# Validates that the host is configured to use the redundant pfv-netinfra-01/02
# pair for name resolution and time, and that both servers actually answer.
set -euo pipefail
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
# The authoritative pair (pfv-netinfra-01 / pfv-netinfra-02).
DNS_PRIMARY="192.168.3.252"
DNS_SECONDARY="192.168.3.253"
NTP_PRIMARY="192.168.3.252"
NTP_SECONDARY="192.168.3.253"
RESOLV_CONF="/etc/resolv.conf"
NTP_CONF="/etc/ntpsec/ntp.conf"
# A name every recursive resolver must be able to resolve.
DNS_PROBE_NAME="github.com"
failed=0
have() { command -v "$1" >/dev/null 2>&1; }
# --- Configuration assertions -------------------------------------------------
function test_dns_config_present() {
echo "🔍 Checking $RESOLV_CONF ..."
local problems=0
if [[ -L "$RESOLV_CONF" ]]; then
echo "$RESOLV_CONF is a symlink (would be overwritten by a resolver manager)"
((++problems))
elif [[ ! -f "$RESOLV_CONF" ]]; then
echo "$RESOLV_CONF missing"
((++problems))
fi
for ns in "$DNS_PRIMARY" "$DNS_SECONDARY"; do
if grep -Eq "^[[:space:]]*nameserver[[:space:]]+$ns" "$RESOLV_CONF" 2>/dev/null; then
echo "✅ nameserver $ns configured"
else
echo "❌ nameserver $ns NOT in $RESOLV_CONF"
((++problems))
fi
done
return $problems
}
function test_ntp_config_present() {
echo "🔍 Checking $NTP_CONF ..."
if [[ ! -f "$NTP_CONF" ]]; then
echo "$NTP_CONF missing (is ntpsec installed?)"
return 1
fi
local problems=0
for s in "$NTP_PRIMARY" "$NTP_SECONDARY"; do
if grep -Eq "^[[:space:]]*(server|pool)[[:space:]]+$s" "$NTP_CONF"; then
echo "✅ NTP server $s configured"
else
echo "❌ NTP server $s NOT in $NTP_CONF"
((++problems))
fi
done
return $problems
}
# --- Functional assertions: each server actually answers ----------------------
function _dns_resolves() {
# $1 = server ip. Returns 0 if it resolves DNS_PROBE_NAME.
local server="$1"
if have dig; then
dig @"$server" +short +time=4 +tries=1 "$DNS_PROBE_NAME" A 2>/dev/null | grep -Eq '[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+'
elif have nslookup; then
nslookup "$DNS_PROBE_NAME" "$server" 2>/dev/null | grep -Eq 'Address:[[:space:]]*[0-9]'
elif have host; then
host "$DNS_PROBE_NAME" "$server" 2>/dev/null | grep -Eq 'has address'
else
# Last resort: the resolver itself.
getent ahostsv4 "$DNS_PROBE_NAME" >/dev/null 2>&1
fi
}
function test_dns_servers_answer() {
echo "🔍 Probing DNS servers ..."
local problems=0
for ns in "$DNS_PRIMARY" "$DNS_SECONDARY"; do
if _dns_resolves "$ns"; then
echo "$ns resolves $DNS_PROBE_NAME"
else
echo "$ns did not resolve $DNS_PROBE_NAME"
((++problems))
fi
done
return $problems
}
function _ntp_answers() {
# $1 = server ip. Returns 0 if it responds to a time query.
local server="$1"
if have ntpdate; then
timeout 8 ntpdate -q "$server" 2>/dev/null | grep -Eq 'no-leap|leap'
elif have sntp; then
timeout 8 sntp -t 4 "$server" >/dev/null 2>&1
elif have chronyc; then
# NTS/chrony not expected here, but be tolerant.
chronyc -n -h "$server" tracking >/dev/null 2>&1
else
return 2 # cannot test
fi
}
function test_ntp_servers_answer() {
echo "🔍 Probing NTP servers ..."
local problems=0
for s in "$NTP_PRIMARY" "$NTP_SECONDARY"; do
if _ntp_answers "$s"; then
echo "✅ NTP $s responds to time query"
else
echo "$s did not respond to NTP query"
((++problems))
fi
done
return $problems
}
# --- End-to-end: the host is actually USING the pair --------------------------
function test_resolver_endtoend() {
echo "🔍 End-to-end resolution via $RESOLV_CONF ..."
if getent ahostsv4 "$DNS_PROBE_NAME" >/dev/null 2>&1; then
echo "✅ Host resolves $DNS_PROBE_NAME via configured resolver"
return 0
else
echo "❌ Host cannot resolve $DNS_PROBE_NAME via configured resolver"
return 1
fi
}
function test_ntp_daemon_peers() {
echo "🔍 NTP daemon peer list ..."
local peers
if have ntpq; then
peers="$(ntpq -pn 2>/dev/null || true)"
elif have chronyc; then
peers="$(chronyc -n sources 2>/dev/null || true)"
else
echo "⚠️ No ntpq/chronyc available; skipping daemon peer check"
return 0
fi
local problems=0
for s in "$NTP_PRIMARY" "$NTP_SECONDARY"; do
if echo "$peers" | grep -Eq "^\\s*${s//./\\.}"; then
echo "✅ NTP daemon has peer $s"
else
echo "❌ NTP daemon is NOT tracking $s"
((++problems))
fi
done
# Sync status is informational only: a freshly started daemon needs several
# polls before the reach counter stabilises, so we warn rather than fail.
if echo "$peers" | grep -Eq '\*'; then
echo "✅ NTP daemon reports a synced peer"
else
echo "⚠️ NTP daemon not yet synced (normal for a few minutes after restart)"
fi
return $problems
}
# --- Main ---------------------------------------------------------------------
function main() {
echo "🛰️ Running Redundant DNS/NTP Validation Tests"
echo "================================================"
local total_failures=0
test_dns_config_present || ((++total_failures))
test_ntp_config_present || ((++total_failures))
test_dns_servers_answer || ((++total_failures))
test_ntp_servers_answer || ((++total_failures))
test_resolver_endtoend || ((++total_failures))
test_ntp_daemon_peers || ((++total_failures))
echo "================================================"
if [[ $total_failures -eq 0 ]]; then
echo "✅ All redundant DNS/NTP validation tests passed"
exit 0
else
echo "$total_failures redundant DNS/NTP tests failed"
exit 1
fi
}
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
@@ -0,0 +1,142 @@
#!/bin/bash
# System Requirements Validation Test
# Validates minimum system requirements before deployment
set -euo pipefail
# Test configuration
MIN_RAM_GB=2
MIN_DISK_GB=10
REQUIRED_COMMANDS=("curl" "wget" "git" "systemctl" "apt-get")
# Test functions
function test_memory_requirements() {
local total_mem_kb=$(grep MemTotal /proc/meminfo | awk '{print $2}')
local total_mem_gb=$((total_mem_kb / 1024 / 1024))
if [[ $total_mem_gb -ge $MIN_RAM_GB ]]; then
echo "✅ Memory requirement met: ${total_mem_gb}GB >= ${MIN_RAM_GB}GB"
return 0
else
echo "❌ Memory requirement not met: ${total_mem_gb}GB < ${MIN_RAM_GB}GB"
return 1
fi
}
function test_disk_space() {
local available_gb=$(df / | tail -1 | awk '{print int($4/1024/1024)}')
if [[ $available_gb -ge $MIN_DISK_GB ]]; then
echo "✅ Disk space requirement met: ${available_gb}GB >= ${MIN_DISK_GB}GB"
return 0
else
echo "❌ Disk space requirement not met: ${available_gb}GB < ${MIN_DISK_GB}GB"
return 1
fi
}
function test_required_commands() {
local failed=0
for cmd in "${REQUIRED_COMMANDS[@]}"; do
if command -v "$cmd" >/dev/null 2>&1; then
echo "✅ Required command available: $cmd"
else
echo "❌ Required command missing: $cmd"
((++failed))
fi
done
return $failed
}
function test_os_compatibility() {
if [[ -f /etc/os-release ]]; then
local os_id=$(grep "^ID=" /etc/os-release | cut -d'=' -f2 | tr -d '"')
local os_version=$(grep "^VERSION_ID=" /etc/os-release | cut -d'=' -f2 | tr -d '"')
case "$os_id" in
ubuntu|debian)
echo "✅ OS compatibility: $os_id $os_version (supported)"
return 0
;;
*)
echo "⚠️ OS compatibility: $os_id $os_version (may work, not fully tested)"
return 0
;;
esac
else
echo "❌ Cannot determine OS version"
return 1
fi
}
function test_network_connectivity() {
local test_urls=(
"https://archive.ubuntu.com"
"https://linux.dell.com"
"https://download.proxmox.com"
"https://github.com"
)
local failed=0
for url in "${test_urls[@]}"; do
if curl -s --connect-timeout 10 --max-time 30 "$url" >/dev/null 2>&1; then
echo "✅ Network connectivity: $url"
else
echo "❌ Network connectivity failed: $url"
((++failed))
fi
done
return $failed
}
function test_permissions() {
local test_dirs=("/etc" "/usr/local/bin" "/var/log")
local failed=0
for dir in "${test_dirs[@]}"; do
if [[ -w "$dir" ]]; then
echo "✅ Write permission: $dir"
else
echo "❌ Write permission denied: $dir"
((++failed))
fi
done
return $failed
}
# Main test execution
function main() {
echo "🔍 Running System Requirements Validation"
echo "========================================"
local total_failures=0
# Run all validation tests
test_memory_requirements || ((total_failures++))
test_disk_space || ((total_failures++))
test_required_commands || ((total_failures++))
test_os_compatibility || ((total_failures++))
test_network_connectivity || ((total_failures++))
test_permissions || ((total_failures++))
echo "========================================"
if [[ $total_failures -eq 0 ]]; then
echo "✅ All system requirements validation tests passed"
exit 0
else
echo "$total_failures system requirements validation tests failed"
exit 1
fi
}
# Run main if executed directly
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
+330
View File
@@ -0,0 +1,330 @@
#!/usr/bin/bash
#
# vm-validation.sh
#
# End-to-end validation driver for KNELServerBuild on a sandbox VM.
#
# This script drives a Proxmox VM through: snapshot -> deploy -> validate, with
# one-command rollback. It is designed to be re-run after code fixes are pushed.
#
# DESIGN: deployment is GIT-BASED. The VM clones (or pulls) the public repo
# itself, exactly as a real fresh server would — so the result is identical no
# matter who runs this script (no reliance on a local working copy or rsync).
# All SSH/SCP access goes through Project-Tests/remote.sh; never call ssh here.
#
# USAGE:
# # Discover the numeric VMID on Proxmox:
# ./Project-Tests/vm-validation.sh find-vmid
#
# # Full loop (snapshot + deploy + validate), auto-rollback on failure:
# VM_ID=6000 ./Project-Tests/vm-validation.sh all
#
# # Individual steps:
# VM_ID=6000 ./Project-Tests/vm-validation.sh snapshot
# VM_ID=6000 ./Project-Tests/vm-validation.sh deploy
# VM_ID=6000 ./Project-Tests/vm-validation.sh validate
# VM_ID=6000 ./Project-Tests/vm-validation.sh rollback [snapshot-name]
#
# # Clean re-deploy from scratch (delete + re-clone on VM):
# VM_ID=6000 CLEAN_CLONE=1 ./Project-Tests/vm-validation.sh deploy
#
# CONFIG (override via env, all have sensible defaults):
# PROX_HOST Proxmox node hostname (default: pfv-tsys5)
# PROX_USER SSH user on Proxmox (default: root)
# VM_NAME VM name for VMID lookup/logging (default: sectestbed-sandbox)
# VM_IP VM IP for SSH (default: 192.168.3.50)
# VM_USER SSH user on the VM (default: localuser)
# VM_ID Numeric VMID on Proxmox (REQUIRED except for find-vmid)
# REPO_URL git URL the VM clones (default: https://git.knownelement.com/KNEL/KNELServerBuild.git)
# REMOTE_REPO clone dir under ~$VM_USER (default: KNELServerBuild)
# SNAP_PREFIX snapshot name prefix (default: pre-knel-deploy)
# CLEAN_CLONE if set, delete + re-clone on VM (default: unset)
#
set -uo pipefail
# ---------------------------------------------------------------------------
# Config
# ---------------------------------------------------------------------------
PROX_HOST="${PROX_HOST:-pfv-tsys5}"
PROX_USER="${PROX_USER:-root}"
VM_NAME="${VM_NAME:-sectestbed-sandbox}"
VM_IP="${VM_IP:-192.168.3.50}"
VM_USER="${VM_USER:-localuser}"
VM_ID="${VM_ID:-}"
REPO_URL="${REPO_URL:-https://git.knownelement.com/KNEL/KNELServerBuild.git}"
REMOTE_REPO="${REMOTE_REPO:-KNELServerBuild}"
SNAP_PREFIX="${SNAP_PREFIX:-pre-knel-deploy}"
ACCESS_PUBKEY="${ACCESS_PUBKEY:-$HOME/.ssh/id_ed25519.pub}"
# Re-inject the validation pubkey after each deploy (secharden-ssh replaces
# authorized_keys with the managed production key set, locking out the
# bootstrap/dev key). Set RESTORE_ACCESS=0 to disable.
RESTORE_ACCESS="${RESTORE_ACCESS:-1}"
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_SRC="$(cd "$HERE/.." && pwd)"
REMOTE="$HERE/remote.sh"
STAMP="$(date +%Y%m%d-%H%M%S)"
SNAP_NAME="${SNAP_PREFIX}-${STAMP}"
LOCAL_LOG_DIR="$REPO_SRC/logs/vm-validation"
mkdir -p "$LOCAL_LOG_DIR"
LOCAL_LOG="$LOCAL_LOG_DIR/run-${STAMP}.log"
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
log() { printf '[%s] %s\n' "$(date +%H:%M:%S)" "$*" | tee -a "$LOCAL_LOG"; }
die() { log "ERROR: $*"; exit 1; }
# All remote access funnels through remote.sh.
vm() { bash "$REMOTE" vm "$@"; } # as $VM_USER (SSH)
vmroot() { bash "$REMOTE" vmroot "$@"; } # as root via sudo (SSH)
vmfile() { bash "$REMOTE" vm-file "$@"; } # run local script on VM (SSH)
vmguest() { bash "$REMOTE" vm-guest "$@"; } # as root via guest agent (no SSH/2FA)
prox() { bash "$REMOTE" prox "$@"; } # as $PROX_USER on Proxmox
require_vm_id() {
[[ -n "$VM_ID" ]] || die "VM_ID is required for this command. Find it with: $0 find-vmid"
}
wait_for_vm_ssh() {
log "Waiting for SSH on ${VM_USER}@${VM_IP} to come up..."
for i in $(seq 1 60); do
if vm 'true' >/dev/null 2>&1; then
log "SSH is up (after ${i} tries)."
return 0
fi
sleep 5
done
die "VM did not become SSH-reachable within 5 minutes."
}
# Resolve the ABSOLUTE path of the repo clone on the VM (as $VM_USER) and echo
# it. Using an absolute path avoids the '~' -> root's home trap under sudo.
resolve_remote_repo() {
local p
p="$(vm "cd ~/${REMOTE_REPO} 2>/dev/null && pwd" 2>/dev/null)"
[[ -n "$p" ]] || p="$(vmguest "cd ~${VM_USER}/${REMOTE_REPO} 2>/dev/null && pwd" 2>/dev/null)"
printf '%s' "$p"
}
# Re-inject the validation pubkey into ~$VM_USER/.ssh/authorized_keys OUT OF
# BAND via the Proxmox guest agent (qm guest exec runs as root inside the VM
# and does not depend on SSH). This is necessary because secharden-ssh replaces
# authorized_keys with the managed production key set, which would otherwise
# lock out the bootstrap key used to drive validation. No-op if SSH still works.
restore_vm_access() {
[[ "$RESTORE_ACCESS" = "1" ]] || { log "RESTORE_ACCESS=0; skipping access restore."; return 0; }
[[ -f "$ACCESS_PUBKEY" ]] || { log "WARN: ACCESS_PUBKEY not found ($ACCESS_PUBKEY); cannot restore access."; return 0; }
if vm 'true' >/dev/null 2>&1; then
log "SSH access already works; no need to restore."
return 0
fi
log "SSH access lost (expected after secharden-ssh). Restoring via Proxmox guest agent..."
local payload_b64
# Leading newline guards against the managed authorized_keys lacking a
# trailing newline (which would otherwise concatenate two keys into one).
payload_b64="$(printf '\n%s' "$(cat "$ACCESS_PUBKEY")" | base64 -w0)"
prox "qm guest exec $VM_ID -- /bin/sh -c 'echo $payload_b64 | base64 -d >> /home/${VM_USER}/.ssh/authorized_keys'" \
>/dev/null 2>&1 || { log "WARN: guest-agent key append failed."; return 0; }
prox "qm guest exec $VM_ID -- /bin/sh -c 'chown ${VM_USER}:${VM_USER} /home/${VM_USER}/.ssh/authorized_keys; chmod 600 /home/${VM_USER}/.ssh/authorized_keys'" \
>/dev/null 2>&1 || true
if vm 'true' >/dev/null 2>&1; then
log "Access restored."
return 0
fi
# If SSH still fails after re-injecting the key, 2FA is almost certainly the
# cause (secharden-2fa enforces publickey+keyboard-interactive, which no
# non-interactive SSH client can satisfy). That is expected and not fatal:
# the guest agent still gives us full out-of-band access for log fetch and
# the validation suite.
if vmguest 'grep -q "^AuthenticationMethods" /etc/ssh/sshd_config' >/dev/null 2>&1; then
log "SSH requires 2FA (expected after secharden-2fa); using guest agent for further access."
else
log "WARN: access still not working after restore and 2FA not detected. Check sshd_config."
fi
}
# ---------------------------------------------------------------------------
# Commands
# ---------------------------------------------------------------------------
cmd_find_vmid() {
log "Listing VMs on Proxmox host '$PROX_HOST' matching '$VM_NAME':"
prox 'qm list' 2>&1 | tee -a "$LOCAL_LOG" \
| { IFS= read -r header; echo "$header"; grep -i "$VM_NAME" || true; }
log "Set VM_ID=<number> env var based on the row above."
}
cmd_snapshot() {
require_vm_id
log "Creating snapshot '$SNAP_NAME' of VMID $VM_ID on $PROX_HOST..."
prox "qm snapshot $VM_ID $SNAP_NAME --vmstate 1" 2>&1 | tee -a "$LOCAL_LOG" \
|| die "Snapshot creation failed."
echo "$SNAP_NAME" > "$LOCAL_LOG_DIR/.last-snapshot"
log "Snapshot '$SNAP_NAME' recorded as rollback target."
}
cmd_rollback() {
require_vm_id
local target="${1:-$(cat "$LOCAL_LOG_DIR/.last-snapshot" 2>/dev/null || true)}"
[[ -n "$target" ]] || die "No snapshot name given and no .last-snapshot on disk."
log "Rolling back VMID $VM_ID to snapshot '$target'..."
# Proxmox rollback requires the VM to be stopped.
prox "qm stop $VM_ID" 2>&1 | tee -a "$LOCAL_LOG" || true
sleep 5
prox "qm rollback $VM_ID $target" 2>&1 | tee -a "$LOCAL_LOG" \
|| die "Rollback command failed."
log "Starting VMID $VM_ID..."
prox "qm start $VM_ID" 2>&1 | tee -a "$LOCAL_LOG" || true
wait_for_vm_ssh
log "Rollback complete."
}
# Ensure the VM has git + ca-certificates (fresh-server bootstrap).
bootstrap_git_on_vm() {
log "Ensuring git is present on the VM..."
vm 'command -v git >/dev/null 2>&1 || sudo -n DEBIAN_FRONTEND=noninteractive apt-get -y -qq install git ca-certificates' \
2>&1 | tee -a "$LOCAL_LOG" || die "Failed to bootstrap git on VM."
vm 'sudo -n DEBIAN_FRONTEND=noninteractive apt-get -y -qq install ca-certificates' 2>&1 | tee -a "$LOCAL_LOG" || true
}
# Clone or pull the repo on the VM. Returns absolute path on stdout (via log).
sync_repo_on_vm() {
bootstrap_git_on_vm
if [[ -n "${CLEAN_CLONE:-}" ]]; then
log "CLEAN_CLONE set: removing existing clone on VM."
vm "rm -rf ~/${REMOTE_REPO}" 2>&1 | tee -a "$LOCAL_LOG" || true
fi
log "Ensuring repo is cloned/pulled on the VM from:"
log " $REPO_URL"
vm "
set -e
if [ -d ~/${REMOTE_REPO}/.git ]; then
cd ~/${REMOTE_REPO}
git fetch --all --prune
git reset --hard origin/HEAD 2>/dev/null || git reset --hard origin/main
git clean -xfd
else
git clone --filter=blob:none '$REPO_URL' ~/${REMOTE_REPO}
cd ~/${REMOTE_REPO}
fi
git log --oneline -1
" 2>&1 | tee -a "$LOCAL_LOG" || die "Repo sync failed on VM."
log "Repo ready on VM."
}
# The remote setup runner: a self-contained script we ship to the VM so the
# sudo'd setup runs from a known-good absolute path with full logging. Using a
# file avoids nested-quote hell across local -> ssh -> sudo -> bash -c.
deploy_runner_script() {
cat <<RUNNER
#!/usr/bin/bash
# remote-setup-runner.sh (generated by vm-validation.sh)
# Runs ProjectCode/SetupNewSystem.sh from the repo given by \$1, as root.
set -uo pipefail
# Ensure a sane TERM so the framework's tput-based color helpers work when run
# over a non-interactive SSH session (which has no TTY/TERM by default).
export TERM="\${TERM:-linux}"
REPO_ABS="\${1:?repo abs path required}"
REMOTE_LOG="/tmp/knel-setup.log"
echo "=== KNEL SetupNewSystem start: \$(date -Is) repo=\$REPO_ABS ===" | tee -a "\$REMOTE_LOG"
cd "\$REPO_ABS/ProjectCode" || { echo "FATAL: ProjectCode missing at \$REPO_ABS"; exit 2; }
bash SetupNewSystem.sh 2>&1 | tee -a "\$REMOTE_LOG"
rc=\${PIPESTATUS[0]}
echo "=== KNEL SetupNewSystem end: rc=\$rc \$(date -Is) ===" | tee -a "\$REMOTE_LOG"
exit \$rc
RUNNER
}
cmd_deploy() {
require_vm_id
sync_repo_on_vm
local repo_abs
repo_abs="$(resolve_remote_repo)"
[[ -n "$repo_abs" ]] || die "Could not resolve absolute repo path on VM."
log "Repo absolute path on VM: $repo_abs"
# Ship the runner script and execute it as root via sudo, passing abs path.
local runner_local="$LOCAL_LOG_DIR/remote-setup-runner.sh"
deploy_runner_script > "$runner_local"
vm "mkdir -p ~/${REMOTE_REPO}/Project-Tests/.run" 2>&1 | tee -a "$LOCAL_LOG"
bash "$REMOTE" vm-copy "$runner_local" "${REMOTE_REPO}/Project-Tests/.run/remote-setup-runner.sh" \
2>&1 | tee -a "$LOCAL_LOG" || die "Failed to ship runner script."
log "Running SetupNewSystem.sh on the VM as root (this takes several minutes)..."
# Resolve abs runner path the same way (no ~ under sudo).
local runner_abs
runner_abs="$(vm "cd ~/${REMOTE_REPO}/Project-Tests/.run && pwd")/remote-setup-runner.sh"
vmroot "bash '$runner_abs' '$repo_abs'" 2>&1 | tee -a "$LOCAL_LOG" || true
# secharden-ssh (run near the end of setup) replaces authorized_keys with the
# managed production key set, locking out the bootstrap key. Restore the
# validation key out-of-band BEFORE we try to fetch the log over SSH.
restore_vm_access
# Fetch the remote log for full fidelity (strip ANSI color codes). SSH works
# only until secharden-2fa flips 2FA on; after that, use the guest agent.
local fetch_cmd="sed -r 's/\\x1B\\[[0-9;]*[mK]//g' /tmp/knel-setup.log 2>/dev/null || cat /tmp/knel-setup.log"
if ! vm "$fetch_cmd" > "$LOCAL_LOG_DIR/setup-output-${STAMP}.log" 2>/dev/null; then
vmguest "$fetch_cmd" > "$LOCAL_LOG_DIR/setup-output-${STAMP}.log" 2>/dev/null || true
fi
# Detect the exit marker. Prefer the full fetched log, but always fall back
# to the live stream ($LOCAL_LOG) which is captured regardless of whether
# post-setup SSH/2FA let us fetch the remote log.
local rc_marker
rc_marker=$(grep -oE 'rc=[0-9]+' "$LOCAL_LOG_DIR/setup-output-${STAMP}.log" 2>/dev/null | tail -1 || true)
[[ -n "$rc_marker" ]] || rc_marker=$(grep -oE 'rc=[0-9]+' "$LOCAL_LOG" 2>/dev/null | tail -1 || true)
log "Setup run finished. Marker: ${rc_marker:-unknown}"
if [[ "${rc_marker:-}" != "rc=0" ]]; then
log "Setup did NOT complete cleanly. See: $LOCAL_LOG_DIR/setup-output-${STAMP}.log (and $LOCAL_LOG)"
return 1
fi
log "Setup completed successfully."
}
cmd_validate() {
require_vm_id
log "Running post-deploy validation suite on the VM..."
local repo_abs
repo_abs="$(resolve_remote_repo)"
[[ -n "$repo_abs" ]] || die "Could not resolve absolute repo path on VM."
# Prefer SSH; fall back to the guest agent (post-2FA SSH needs a TOTP token).
if ! vmroot "cd '$repo_abs' && bash Project-Tests/run-tests.sh all" 2>&1 | tee -a "$LOCAL_LOG"; then
vmguest "cd '$repo_abs' && bash Project-Tests/run-tests.sh all" 2>&1 | tee -a "$LOCAL_LOG" || true
fi
log "Validation run finished. Inspect output above / in $LOCAL_LOG."
}
cmd_all() {
require_vm_id
log "=== FULL VALIDATION LOOP: $VM_NAME (VMID $VM_ID) ==="
cmd_snapshot
if cmd_deploy && cmd_validate; then
log "=== ALL GREEN ==="
return 0
fi
log "=== FAILURE — auto-rolling back to '$SNAP_NAME' ==="
cmd_rollback "$SNAP_NAME"
log "Rolled back. Fix and push, then re-run: VM_ID=$VM_ID $0 deploy && VM_ID=$VM_ID $0 validate"
return 1
}
# ---------------------------------------------------------------------------
# Dispatch
# ---------------------------------------------------------------------------
subcmd="${1:-}"
case "$subcmd" in
find-vmid) cmd_find_vmid ;;
snapshot) cmd_snapshot ;;
deploy) cmd_deploy ;;
validate) cmd_validate ;;
rollback) cmd_rollback "${2:-}" ;;
all) cmd_all ;;
""|-h|--help|help)
sed -n '2,49p' "${BASH_SOURCE[0]}" >&2
exit 0
;;
*) die "Unknown command '$subcmd'. Run '$0 help'." ;;
esac
@@ -0,0 +1,2 @@
#/etc/cockpit/disallowed-users
# List of users which are not allowed to login to Cockpit
@@ -0,0 +1,14 @@
option rfc3442-classless-static-routes code 121 = array of unsigned integer 8;
send host-name = gethostname();
request subnet-mask, broadcast-address, time-offset, routers,
domain-name, host-name,
domain-name-servers, domain-search, ntp-servers,
rfc3442-classless-static-routes;
# Pin DNS and NTP to the redundant pfv-netinfra-01/02 pair regardless of what
# the DHCP server advertises, so every host on this build uses the same
# authoritative recursive resolvers and time sources.
supersede domain-name-servers 192.168.3.252, 192.168.3.253;
supersede domain-search "knel.net";
supersede ntp-servers 192.168.3.252, 192.168.3.253;
@@ -0,0 +1 @@
install cramfs /bin/true
@@ -0,0 +1 @@
install dccp /bin/true
@@ -0,0 +1 @@
install freevxfs /bin/true
@@ -0,0 +1 @@
install hfs /bin/true
@@ -0,0 +1 @@
install hfsplus /bin/true
@@ -0,0 +1 @@
install jffs2 /bin/true
@@ -0,0 +1 @@
install rds /bin/true
@@ -0,0 +1 @@
install sctp /bin/true
@@ -0,0 +1 @@
install squashfs /bin/true
@@ -0,0 +1 @@
install tipc /bin/true
@@ -0,0 +1 @@
install udf /bin/true
@@ -0,0 +1 @@
install usb-storage /bin/true
@@ -0,0 +1,21 @@
driftfile /var/lib/ntp/ntp.drift
leapfile /usr/share/zoneinfo/leap-seconds.list
# Redundant upstream time sources: pfv-netinfra-01/02 (Technitium/Pi-hole hosts
# also serving NTP). IPs are used (not hostnames) because the knel.net name for
# these hosts resolves to a Tailscale CGNAT address, not the LAN address, and
# because NTP must come up before DNS is available. iburst speeds initial sync.
server 192.168.3.252 iburst
server 192.168.3.253 iburst
# Hardened client: sync from the configured servers but never serve time to
# anyone else. Note: `interface listen 127.0.0.1` must NOT be used here — it
# binds ntpd to loopback, making outbound queries carry a 127.0.0.1 source
# address that upstream servers cannot reply to (symptoms: peers stuck in
# .INIT. with reach 0). Use restrict rules to control access instead.
restrict default ignore
restrict 127.0.0.1
restrict ::1
restrict 192.168.3.252 nomodify notrap nopeer
restrict 192.168.3.253 nomodify notrap nopeer
@@ -0,0 +1,11 @@
# Managed by KNELServerBuild — do not edit; changes will be overwritten.
#
# Redundant recursive DNS via pfv-netinfra-01/02 (Technitium + Pi-hole).
# IPs are used (required: nameserver directives must be addresses, and the
# knel.net name for these hosts resolves to a Tailscale CGNAT address rather
# than the LAN address). If the primary is unreachable, glibc's resolver
# automatically falls through to the secondary.
domain knel.net
search knel.net
nameserver 192.168.3.252
nameserver 192.168.3.253
@@ -0,0 +1 @@
/.*/ tsysrootaccount@knel.net
@@ -0,0 +1,34 @@
#!/bin/bash
#curl -s http://dl.turnsys.net/omsa.sh|/bin/bash
gpg --keyserver hkp://pool.sks-keyservers.net:80 --recv-key 1285491434D8786F
gpg -a --export 1285491434D8786F | apt-key add -
echo "deb https://linux.dell.com/repo/community/openmanage/930/bionic bionic main" > /etc/apt/sources.list.d/linux.dell.com.sources.list
wget https://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman-curl-client-transport1_2.6.5-0ubuntu3_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman-client4_2.6.5-0ubuntu3_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman1_2.6.5-0ubuntu3_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/libwsman-server1_2.6.5-0ubuntu3_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/universe/s/sblim-sfcc/libcimcclient0_2.2.8-0ubuntu2_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/universe/o/openwsman/openwsman_2.6.5-0ubuntu3_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/multiverse/c/cim-schema/cim-schema_2.48.0-0ubuntu1_all.deb
wget https://archive.ubuntu.com/ubuntu/pool/universe/s/sblim-sfc-common/libsfcutil0_1.0.1-0ubuntu4_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/multiverse/s/sblim-sfcb/sfcb_1.4.9-0ubuntu5_amd64.deb
wget https://archive.ubuntu.com/ubuntu/pool/universe/s/sblim-cmpi-devel/libcmpicppimpl0_2.0.3-0ubuntu2_amd64.deb
dpkg -i libwsman-curl-client-transport1_2.6.5-0ubuntu3_amd64.deb
dpkg -i libwsman-client4_2.6.5-0ubuntu3_amd64.deb
dpkg -i libwsman1_2.6.5-0ubuntu3_amd64.deb
dpkg -i libwsman-server1_2.6.5-0ubuntu3_amd64.deb
dpkg -i libcimcclient0_2.2.8-0ubuntu2_amd64.deb
dpkg -i openwsman_2.6.5-0ubuntu3_amd64.deb
dpkg -i cim-schema_2.48.0-0ubuntu1_all.deb
dpkg -i libsfcutil0_1.0.1-0ubuntu4_amd64.deb
dpkg -i sfcb_1.4.9-0ubuntu5_amd64.deb
dpkg -i libcmpicppimpl0_2.0.3-0ubuntu2_amd64.deb
apt update
apt -y install srvadmin-all
touch /opt/dell/srvadmin/lib64/openmanage/IGNORE_GENERATION
#logout,login, then run
# srvadmin-services.sh enable && srvadmin-services.sh start
@@ -8,13 +8,13 @@ OPENSSL_FILE="openssl-1.1.0h.tar.gz"
NGHTTP_URL_BASE="https://github.com/nghttp2/nghttp2/releases/download/v1.31.0/" NGHTTP_URL_BASE="https://github.com/nghttp2/nghttp2/releases/download/v1.31.0/"
NGHTTP_FILE="nghttp2-1.31.0.tar.gz" NGHTTP_FILE="nghttp2-1.31.0.tar.gz"
APR_URL_BASE="http://mirrors.whoishostingthis.com/apache/apr/" APR_URL_BASE="https://archive.apache.org/dist/apr/"
APR_FILE="apr-1.6.3.tar.gz" APR_FILE="apr-1.6.3.tar.gz"
APR_UTIL_URL_BASE="http://mirrors.whoishostingthis.com/apache/apr/" APR_UTIL_URL_BASE="https://archive.apache.org/dist/apr/"
APR_UTIL_FILE="apr-util-1.6.1.tar.gz" APR_UTIL_FILE="apr-util-1.6.1.tar.gz"
APACHE_URL_BASE="http://mirrors.whoishostingthis.com/apache/httpd/" APACHE_URL_BASE="https://archive.apache.org/dist/httpd/"
APACHE_FILE="httpd-2.4.33.tar.gz" APACHE_FILE="httpd-2.4.33.tar.gz"
CURL_URL_BASE="https://curl.haxx.se/download/" CURL_URL_BASE="https://curl.haxx.se/download/"
@@ -0,0 +1,426 @@
#!/bin/bash
# TSYS Security Hardening - Two-Factor Authentication
# Implements 2FA for SSH, Cockpit, and Webmin services
# Uses Google Authenticator (TOTP) for time-based tokens
#####
#Core framework functions...
#####
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
export PROJECT_ROOT_PATH
PROJECT_ROOT_PATH="$(cd "$SCRIPT_DIR/../../.." && pwd)"
export GIT_VENDOR_PATH_ROOT
GIT_VENDOR_PATH_ROOT="$PROJECT_ROOT_PATH/vendor/git@git.knownelement.com/29418/"
export KNELShellFrameworkRoot
KNELShellFrameworkRoot="$GIT_VENDOR_PATH_ROOT/KNEL/KNELShellFramework"
source "$KNELShellFrameworkRoot/Framework-ConfigFiles/FrameworkVars"
for framework_include_file in "$KNELShellFrameworkRoot"/Framework-Includes/*; do
source "$framework_include_file"
done
for project_include_file in "$PROJECT_ROOT_PATH"/Project-Includes/*; do
source "$project_include_file"
done
# 2FA Configuration
BACKUP_DIR="/root/backup/2fa"
PAM_CONFIG_DIR="/etc/pam.d"
SSH_CONFIG="/etc/ssh/sshd_config"
COCKPIT_CONFIG="/etc/cockpit/cockpit.conf"
# Create backup directory
mkdir -p "$BACKUP_DIR"
print_info "TSYS Two-Factor Authentication Setup"
# Backup existing configurations
function backup_configs() {
print_info "Creating backup of existing configurations..."
# Backup SSH configuration
if [[ -f "$SSH_CONFIG" ]]; then
cp "$SSH_CONFIG" "$BACKUP_DIR/sshd_config.bak"
print_info "SSH config backed up"
fi
# Backup PAM configurations
if [[ -d "$PAM_CONFIG_DIR" ]]; then
cp -r "$PAM_CONFIG_DIR" "$BACKUP_DIR/pam.d.bak"
print_info "PAM configs backed up"
fi
# Backup Cockpit configuration if exists
if [[ -f "$COCKPIT_CONFIG" ]]; then
cp "$COCKPIT_CONFIG" "$BACKUP_DIR/cockpit.conf.bak"
print_info "Cockpit config backed up"
fi
print_info "Backup completed: $BACKUP_DIR"
}
# Install required packages
function install_2fa_packages() {
print_info "Installing 2FA packages..."
# Update package cache
apt-get update
# Install Google Authenticator PAM module
# Install QR code generator for terminal display
apt-get install -y libpam-google-authenticator qrencode
print_info "2FA packages installed successfully"
}
# Configure SSH for 2FA
function configure_ssh_2fa() {
print_info "Configuring SSH for 2FA..."
# Configure SSH daemon
print_info "Updating SSH configuration..."
# Enable challenge-response authentication
if ! grep -q "^ChallengeResponseAuthentication yes" "$SSH_CONFIG"; then
sed -i 's/^ChallengeResponseAuthentication.*/ChallengeResponseAuthentication yes/' "$SSH_CONFIG" || \
echo "ChallengeResponseAuthentication yes" >> "$SSH_CONFIG"
fi
if ! grep -q "^KbdInteractiveAuthentication yes" "$SSH_CONFIG"; then
sed -i 's/^KbdInteractiveAuthentication.*/KbdInteractiveAuthentication yes/' "$SSH_CONFIG" || \
echo "KbdInteractiveAuthentication yes" >> "$SSH_CONFIG"
fi
# Enable PAM authentication
if ! grep -q "^UsePAM yes" "$SSH_CONFIG"; then
sed -i 's/^UsePAM.*/UsePAM yes/' "$SSH_CONFIG" || \
echo "UsePAM yes" >> "$SSH_CONFIG"
fi
# Configure authentication methods (key + 2FA)
if ! grep -q "^AuthenticationMethods" "$SSH_CONFIG"; then
echo "AuthenticationMethods publickey,keyboard-interactive" >> "$SSH_CONFIG"
else
sed -i 's/^AuthenticationMethods.*/AuthenticationMethods publickey,keyboard-interactive/' "$SSH_CONFIG"
fi
print_info "SSH configuration updated"
}
# Configure PAM for 2FA
function configure_pam_2fa() {
print_info "Configuring PAM for 2FA..."
# Create backup of original PAM SSH config
cp "$PAM_CONFIG_DIR/sshd" "$PAM_CONFIG_DIR/sshd.bak.$(date +%Y%m%d)"
# Configure PAM to use Google Authenticator
cat > "$PAM_CONFIG_DIR/sshd" << 'EOF'
# PAM configuration for SSH with 2FA
# Standard Un*x authentication
@include common-auth
# Google Authenticator 2FA
auth required pam_google_authenticator.so nullok
# Standard Un*x authorization
@include common-account
# SELinux needs to be the first session rule
session required pam_selinux.so close
session required pam_loginuid.so
# Standard Un*x session setup and teardown
@include common-session
# Print the message of the day upon successful login
session optional pam_motd.so motd=/run/motd.dynamic
session optional pam_motd.so noupdate
# Print the status of the user's mailbox upon successful login
session optional pam_mail.so standard noenv
# Set up user limits from /etc/security/limits.conf
session required pam_limits.so
# SELinux needs to intervene at login time
session required pam_selinux.so open
# Standard Un*x password updating
@include common-password
EOF
print_info "PAM configuration updated for SSH 2FA"
}
# Configure Cockpit for 2FA
function configure_cockpit_2fa() {
print_info "Configuring Cockpit for 2FA..."
# Create Cockpit config directory if it doesn't exist
mkdir -p "$(dirname "$COCKPIT_CONFIG")"
# Configure Cockpit to use PAM with 2FA
cat > "$COCKPIT_CONFIG" << 'EOF'
[WebService]
# Enable 2FA for Cockpit web interface
LoginTitle = TSYS Server Management
LoginTo = 300
RequireHost = true
[Session]
# Use PAM for authentication (includes 2FA)
Banner = /etc/cockpit/issue.cockpit
IdleTimeout = 15
EOF
# Create PAM configuration for Cockpit
cat > "$PAM_CONFIG_DIR/cockpit" << 'EOF'
# PAM configuration for Cockpit with 2FA
auth requisite pam_nologin.so
auth required pam_env.so
auth required pam_faillock.so preauth
auth sufficient pam_unix.so try_first_pass
auth required pam_google_authenticator.so nullok
auth required pam_faillock.so authfail
auth required pam_deny.so
account required pam_nologin.so
account include system-auth
account required pam_faillock.so
session required pam_selinux.so close
session required pam_loginuid.so
session optional pam_keyinit.so force revoke
session include system-auth
session required pam_selinux.so open
session optional pam_motd.so
EOF
print_info "Cockpit 2FA configuration completed"
}
# Configure Webmin for 2FA (if installed)
function configure_webmin_2fa() {
print_info "Checking for Webmin installation..."
local webmin_config="/etc/webmin/miniserv.conf"
if [[ -f "$webmin_config" ]]; then
print_info "Webmin found, configuring 2FA..."
# Stop webmin service
systemctl stop webmin || true
# Enable 2FA in Webmin configuration. `sed -i ... || echo` would never
# append, because sed returns 0 even when it matches nothing; guard with
# grep so the directive is added when absent and updated when present.
if grep -q '^twofactor_provider=' "$webmin_config"; then
sed -i 's/^twofactor_provider=.*/twofactor_provider=totp/' "$webmin_config"
else
echo "twofactor_provider=totp" >> "$webmin_config"
fi
# Enable 2FA requirement
if grep -q '^twofactor=' "$webmin_config"; then
sed -i 's/^twofactor=.*/twofactor=1/' "$webmin_config"
else
echo "twofactor=1" >> "$webmin_config"
fi
# Start webmin service
systemctl start webmin || true
print_info "Webmin 2FA configuration completed"
else
print_info "Webmin not found, skipping configuration"
fi
}
# Setup 2FA for users
function setup_user_2fa() {
print_info "Setting up 2FA for system users..."
local users=("localuser" "root")
for user in "${users[@]}"; do
if id "$user" &>/dev/null; then
print_info "Setting up 2FA for user: $user"
local user_home
user_home="$(getent passwd "$user" | cut -d: -f6)"
if [[ -z "$user_home" ]]; then
print_info "No home directory for $user, skipping"
continue
fi
# Create 2FA setup script for user
cat > "/tmp/setup-2fa-$user.sh" << 'EOF'
#!/bin/bash
echo "Setting up Google Authenticator for user: $USER"
echo "Please follow the prompts to configure 2FA:"
echo "1. Answer 'y' to update your time-based token"
echo "2. Scan the QR code with your authenticator app"
echo "3. Save the backup codes in a secure location"
echo "4. Answer 'y' to the remaining questions for security"
echo ""
google-authenticator -t -d -f -r 3 -R 30 -W
EOF
chmod +x "/tmp/setup-2fa-$user.sh"
# Instructions for user setup
cat > "$user_home/2fa-setup-instructions.txt" << EOF
TSYS Two-Factor Authentication Setup Instructions
==============================================
Your system has been configured for 2FA. To complete setup:
1. Install an authenticator app on your phone:
- Google Authenticator
- Authy
- Microsoft Authenticator
2. Run the setup command:
sudo /tmp/setup-2fa-$user.sh
3. Follow the prompts:
- Scan the QR code with your app
- Save the backup codes securely
- Answer 'y' to security questions
4. Test your setup:
- SSH to the server
- Enter your 6-digit code when prompted
IMPORTANT: Save backup codes in a secure location!
Without them, you may be locked out if you lose your phone.
For support, contact your system administrator.
EOF
chown "$user:$user" "$user_home/2fa-setup-instructions.txt"
print_info "2FA setup prepared for user: $user"
else
print_info "User $user not found, skipping"
fi
done
}
# Restart services
function restart_services() {
print_info "Restarting services..."
# Test SSH configuration
if sshd -t; then
systemctl restart sshd
print_info "SSH service restarted"
else
print_error "SSH configuration test failed"
return 1
fi
# Restart Cockpit if installed
if systemctl is-enabled cockpit.socket &>/dev/null; then
systemctl restart cockpit.socket
print_info "Cockpit service restarted"
fi
# Restart Webmin if installed
if systemctl is-enabled webmin &>/dev/null; then
systemctl restart webmin
print_info "Webmin service restarted"
fi
}
# Validation and testing
function validate_2fa_setup() {
print_info "Validating 2FA setup..."
# Check if Google Authenticator is installed
if command -v google-authenticator &>/dev/null; then
print_info "Google Authenticator installed"
else
print_error "Google Authenticator not found"
return 1
fi
# Check SSH configuration
if grep -q "AuthenticationMethods publickey,keyboard-interactive" "$SSH_CONFIG"; then
print_info "SSH 2FA configuration valid"
else
print_error "SSH 2FA configuration invalid"
return 1
fi
# Check PAM configuration
if grep -q "pam_google_authenticator.so" "$PAM_CONFIG_DIR/sshd"; then
print_info "PAM 2FA configuration valid"
else
print_error "PAM 2FA configuration invalid"
return 1
fi
# Check service status
if systemctl is-active sshd &>/dev/null; then
print_info "SSH service is running"
else
print_error "SSH service is not running"
return 1
fi
print_info "2FA validation completed successfully"
}
# Display final instructions
function show_final_instructions() {
print_info "2FA Setup Completed"
print_info "Two-Factor Authentication has been configured for:"
print_info "- SSH (requires key + 2FA token)"
print_info "- Cockpit web interface"
if [[ -f "/etc/webmin/miniserv.conf" ]]; then
print_info "- Webmin administration panel"
fi
print_info "IMPORTANT: Complete user setup immediately!"
print_info "1. Check /home/*/2fa-setup-instructions.txt for user setup"
print_info "2. Run setup scripts for each user"
print_info "3. Test 2FA before logging out"
print_info "Backup location: $BACKUP_DIR"
print_info "To disable 2FA, restore configurations from backup"
print_info "2FA setup completed successfully!"
}
# Main execution
function main() {
# Check if running as root
if [[ $EUID -ne 0 ]]; then
print_error "This script must be run as root"
exit 1
fi
# Execute setup steps
backup_configs
install_2fa_packages
configure_ssh_2fa
configure_pam_2fa
configure_cockpit_2fa
configure_webmin_2fa
setup_user_2fa
restart_services
validate_2fa_setup
show_final_instructions
}
# Run main function
main "$@"
@@ -0,0 +1,50 @@
#!/bin/bash
#####
#Core framework functions...
#####
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
export PROJECT_ROOT_PATH
PROJECT_ROOT_PATH="$(cd "$SCRIPT_DIR/../../.." && pwd)"
export GIT_VENDOR_PATH_ROOT
GIT_VENDOR_PATH_ROOT="$PROJECT_ROOT_PATH/vendor/git@git.knownelement.com/29418/"
export KNELShellFrameworkRoot
KNELShellFrameworkRoot="$GIT_VENDOR_PATH_ROOT/KNEL/KNELShellFramework"
export CONFIGFILES_PATH
CONFIGFILES_PATH="$PROJECT_ROOT_PATH/ProjectCode/ConfigFiles"
source "$KNELShellFrameworkRoot/Framework-ConfigFiles/FrameworkVars"
for framework_include_file in "$KNELShellFrameworkRoot"/Framework-Includes/*; do
source "$framework_include_file"
done
for project_include_file in "$PROJECT_ROOT_PATH"/Project-Includes/*; do
source "$project_include_file"
done
# Material herein Sourced from
# https://cisofy.com/documentation/lynis/
# https://jbcsec.com/configure-linux-ssh/
# https://opensource.com/article/20/5/linux-security-lynis
# https://forum.greenbone.net/t/ssh-authentication/13536
# openvas
#lynis
#Auditd
cat "$CONFIGFILES_PATH/AuditD/auditd.conf" > /etc/audit/auditd.conf
# Systemd
cat "$CONFIGFILES_PATH/Systemd/journald.conf" > /etc/systemd/journald.conf
# logrotate
cat "$CONFIGFILES_PATH/Logrotate/logrotate.conf" > /etc/logrotate.conf
@@ -0,0 +1,126 @@
#!/bin/bash
#########################################
#Core framework functions...
#########################################
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
export PROJECT_ROOT_PATH
PROJECT_ROOT_PATH="$(cd "$SCRIPT_DIR/../../.." && pwd)"
export GIT_VENDOR_PATH_ROOT
GIT_VENDOR_PATH_ROOT="$PROJECT_ROOT_PATH/vendor/git@git.knownelement.com/29418/"
export KNELShellFrameworkRoot
KNELShellFrameworkRoot="$GIT_VENDOR_PATH_ROOT/KNEL/KNELShellFramework"
export CONFIGFILES_PATH
CONFIGFILES_PATH="$PROJECT_ROOT_PATH/ProjectCode/ConfigFiles"
source "$KNELShellFrameworkRoot/Framework-ConfigFiles/FrameworkVars"
for framework_include_file in "$KNELShellFrameworkRoot"/Framework-Includes/*; do
source "$framework_include_file"
done
for project_include_file in "$PROJECT_ROOT_PATH"/Project-Includes/*; do
source "$project_include_file"
done
#########################################
# Core script code begins here
#########################################
# Sourced from
# https://complianceascode.readthedocs.io/en/latest/manual/developer/01_introduction.html
# https://github.com/ComplianceAsCode/content
# https://github.com/ComplianceAsCode
#apparmor
#enforcing
#enabled in bootloader config
#aide
#auditd
#disable auto mounting
#disable usb storage
#motd
#remote login warning banner
#Ensure time sync is working
#systemd-timesync
#ntp
#chrony
#password complexity
#password expiration warning
#password expiration time
#password hashing algo
#fix grub perms
if [ "$IS_RASPI" = 0 ] ; then
chown root:root /boot/grub/grub.cfg
chmod og-rwx /boot/grub/grub.cfg
chmod 0400 /boot/grub/grub.cfg
fi
#disable auto mounting
systemctl --now disable autofs || true
apt-get -y --purge remove autofs || true
#disable usb storage
cat "$CONFIGFILES_PATH/ModProbe/usb_storage.conf" > /etc/modprobe.d/usb_storage.conf
cat "$CONFIGFILES_PATH/ModProbe/dccp.conf" > /etc/modprobe.d/dccp.conf
cat "$CONFIGFILES_PATH/ModProbe/rds.conf" > /etc/modprobe.d/rds.conf
cat "$CONFIGFILES_PATH/ModProbe/sctp.conf" > /etc/modprobe.d/sctp.conf
cat "$CONFIGFILES_PATH/ModProbe/tipc.conf" > /etc/modprobe.d/tipc.conf
cat "$CONFIGFILES_PATH/ModProbe/cramfs.conf" > /etc/modprobe.d/cramfs.conf
cat "$CONFIGFILES_PATH/ModProbe/freevxfs.conf" > /etc/modprobe.d/freevxfs.conf
cat "$CONFIGFILES_PATH/ModProbe/hfs.conf" > /etc/modprobe.d/hfs.conf
cat "$CONFIGFILES_PATH/ModProbe/hfsplus.conf" > /etc/modprobe.d/hfsplus.conf
cat "$CONFIGFILES_PATH/ModProbe/jffs2.conf" > /etc/modprobe.d/jffs2.conf
cat "$CONFIGFILES_PATH/ModProbe/squashfs.conf" > /etc/modprobe.d/squashfs.conf
cat "$CONFIGFILES_PATH/ModProbe/udf.conf" > /etc/modprobe.d/udf.conf
#banners
cat "$CONFIGFILES_PATH/BANNERS/issue" > /etc/issue
cat "$CONFIGFILES_PATH/BANNERS/issue.net" > /etc/issue.net
cat "$CONFIGFILES_PATH/BANNERS/motd" > /etc/motd
#Cron perms
if [ -f /etc/cron.deny ]; then
rm /etc/cron.deny || true
fi
touch /etc/cron.allow
chmod g-wx,o-rwx /etc/cron.allow
chown root:root /etc/cron.allow
chmod og-rwx /etc/crontab
chmod og-rwx /etc/cron.hourly/
chmod og-rwx /etc/cron.daily/
chmod og-rwx /etc/cron.weekly/
chmod og-rwx /etc/cron.monthly/
chown root:root /etc/cron.d/
chmod og-rwx /etc/cron.d/
# At perms
rm -f /etc/at.deny || true
touch /etc/at.allow
chmod g-wx,o-rwx /etc/at.allow
chown root:root /etc/at.allow

Some files were not shown because too many files have changed in this diff Show More