mrcharles 3d69a7a989 feat(perf): census covers cpu + disk, report-only [#757]
CPU via the same docker stats pass joined to app fqdns; disk = root fs
pct + docker system df + top-15 appdata dirs joined to fqdns. CPU/disk
are REPORT-ONLY per t/317 (human-required class).

Details: https://projects.knownelement.com/issues/757#note-1
2026-09-03 11:52:58 -05:00

Knowledge docs live on Discourse — perf write-up: t/316 · VP TechOps category · Work tracking: Redmine — Cloudron project · Infra audit log: t/298

KNEL/cloudron — tsys-cloudron fleet-machine repo

Production Cloudron host my.knownelement.com (tailnet name tsys-cloudron.knel.net). ~55 apps behind the Cloudron panel, platform Postgres/MySQL/mail containers, monitored by Uptime Kuma (Cloudron group).

Path What
perf/ #685 perf work: memory census, pg role-cap SQL, applied sysctl, applied daemon.json, apply-memory-shrink.sh
monitoring/ Uptime Kuma Cloudron app coverage — 55-app inventory + sync checker (moved from PFVCluster, #435/#685/#727)
ansible/ Reserved — AWX/Ansible IaC convergence (#454); playbooks land here as manual changes are codified
scripts/ Rule engine + git hooks (enforcement layer, ADOPTING pattern)
tests/ shellcheck wrapper (Docker — host stays clean)

Ground truth (2026-09-02, #685 — narrative + measurements on t/316)

  • Platform Postgres slot exhaustion root-caused to 7 booking instances + Penpot pinning ~61 idle connections each. Role caps applied (booking 10, Penpot 15, every other app role 20), max_connections 500→700, shared_buffers 2GB, idle-in-transaction timeout 15min. Victim too many clients FATALs → zero.
  • Host sysctl: vm.swappiness=10 (perf/sysctl/). THP already never, virtio scheduler none — optimal, left alone.
  • /etc/docker/daemon.json APPLIED 2026-09-02 (#731): live-restore: true (future dockerd restarts are zero-downtime) + default-ulimits nofile 262144 (every container start gets real fd limits — the postgres prlimit bridge is no longer needed; perf/docker-daemon.json is the SoR copy). NOTE: this host's dockerd runs --log-driver=journald as a unit flag — daemon.json must NOT set log-driver/log-opts or dockerd refuses to start (bit us once, see #731). Container log rotation is journald's job here.
  • Memory shrink table APPLIED 2026-09-02 (#731) via the panel API (perf/apply-memory-shrink.sh): 11 apps resliders, explicit limits 53.4 GB → 29.0 GB (~24 GB of the ~28 GB headroom reclaimed). Skips: photos (manifest floor 3.5 GB), bi.knownelement.com + platform services (do NOT shrink). All 159 apps verified running post-change. 2026-09-03 addendum: OOM-driven bumps (thefnf 512M→1G, grafana 512M→1G, projects 600M→1.5G) + inotify-exhaustion dockerd/containerd incident — see #731 + t/298.

Access

SSH root@my.knownelement.com — routed through the KNEL/PFVCluster tests/remote.sh chokepoint (see AGENTS.md). Panel API token issued 2026-09-02 (~/.creds/cloudron.env, refs only in this repo) — prefer the API over SSH for reads/limits; base image rebuilds (#710) still need SSH.

  • Inotify headroom APPLIED 2026-09-03 (#755, founder-approved): fs.inotify.max_user_instances 128→1024 + max_user_watches →1048576 in the same sysctl file, after containerd/dockerd died of inotify exhaustion (full fleet resync 09:1309:51). Runtime-applied, no daemon restarts; error stream dry from 10:08. OOM fail set complete: thefnf/grafana/projects bumped (1G/1G/1.5G), zero OOMs post-recovery.
  • Swap 4G→8G APPLIED 2026-09-03 (#755, founder-approved): second swapfile /apps.swap2 (4G, fstab-persisted, same-priority family) alongside the original /apps.swap — no swapoff, no restart, live free -h shows 8.0Gi. Idempotent applier: perf/swap/grow-swap2.sh. Headroom for the Monday app push.
S
Description
tsys-cloudron (my.knownelement.com) fleet machine repo — perf tuning, config-as-code, AWX/Ansible IaC [#685]
Readme AGPL-3.0
95 KiB
Languages
Shell 89.3%
Python 10.7%