> Knowledge docs live on Discourse — perf write-up: [t/316](https://community.turnsys.com/t/316) · > [VP TechOps category](https://community.turnsys.com/c/vp-techops) · > Work tracking: [Redmine — Cloudron project](https://projects.knownelement.com/projects/88) · > Infra audit log: [t/298](https://community.turnsys.com/t/298) # KNEL/cloudron — tsys-cloudron fleet-machine repo Production Cloudron host **my.knownelement.com** (tailnet name `tsys-cloudron.knel.net`). ~55 apps behind the Cloudron panel, platform Postgres/MySQL/mail containers, monitored by Uptime Kuma (Cloudron group). | Path | What | |------|------| | `perf/` | [#685](https://projects.knownelement.com/issues/685) perf work: memory census, pg role-cap SQL, applied sysctl, applied `daemon.json`, `apply-memory-shrink.sh` | | `monitoring/` | Uptime Kuma Cloudron app coverage — 55-app inventory + sync checker (moved from PFVCluster, #435/#685/#727) | | `ansible/` | Reserved — AWX/Ansible IaC convergence ([#454](https://projects.knownelement.com/issues/454)); playbooks land here as manual changes are codified | | `scripts/` | Rule engine + git hooks (enforcement layer, ADOPTING pattern) | | `tests/` | shellcheck wrapper (Docker — host stays clean) | ## Ground truth (2026-09-02, #685 — narrative + measurements on [t/316](https://community.turnsys.com/t/316)) - Platform Postgres slot exhaustion root-caused to 7 booking instances + Penpot pinning ~61 idle connections each. Role caps applied (booking 10, Penpot 15, every other app role 20), `max_connections` 500→700, `shared_buffers` 2GB, idle-in-transaction timeout 15min. Victim `too many clients` FATALs → zero. - Host sysctl: `vm.swappiness=10` (`perf/sysctl/`). THP already `never`, virtio scheduler `none` — optimal, left alone. - **`/etc/docker/daemon.json` APPLIED 2026-09-02** (#731): `live-restore: true` (future dockerd restarts are zero-downtime) + `default-ulimits nofile 262144` (every container start gets real fd limits — the postgres prlimit bridge is no longer needed; `perf/docker-daemon.json` is the SoR copy). NOTE: this host's dockerd runs `--log-driver=journald` as a unit flag — daemon.json must NOT set log-driver/log-opts or dockerd refuses to start (bit us once, see #731). Container log rotation is journald's job here. - **Memory shrink table APPLIED 2026-09-02** (#731) via the panel API (`perf/apply-memory-shrink.sh`): 11 apps resliders, explicit limits 53.4 GB → 29.0 GB (~24 GB of the ~28 GB headroom reclaimed). Skips: photos (manifest floor 3.5 GB), `bi.knownelement.com` + platform services (do NOT shrink). All 159 apps verified running post-change. **2026-09-03 addendum:** OOM-driven bumps (thefnf 512M→1G, grafana 512M→1G, projects 600M→1.5G) + inotify-exhaustion dockerd/containerd incident — see #731 + [t/298](https://community.turnsys.com/t/298). ## Access SSH `root@my.knownelement.com` — routed through the KNEL/PFVCluster `tests/remote.sh` chokepoint (see AGENTS.md). Panel API token issued 2026-09-02 (`~/.creds/cloudron.env`, refs only in this repo) — prefer the API over SSH for reads/limits; base image rebuilds (#710) still need SSH. - **Inotify headroom APPLIED 2026-09-03** (#755, founder-approved): `fs.inotify.max_user_instances` 128→1024 + `max_user_watches` →1048576 in the same sysctl file, after containerd/dockerd died of inotify exhaustion (full fleet resync 09:13–09:51). Runtime-applied, no daemon restarts; error stream dry from 10:08. OOM fail set complete: thefnf/grafana/projects bumped (1G/1G/1.5G), zero OOMs post-recovery. - **Swap 4G→8G APPLIED 2026-09-03** (#755, founder-approved): second swapfile `/apps.swap2` (4G, fstab-persisted, same-priority family) alongside the original `/apps.swap` — no swapoff, no restart, live `free -h` shows 8.0Gi. Idempotent applier: `perf/swap/grow-swap2.sh`. Headroom for the Monday app push.