Applied on my.knownelement.com 2026-09-03 (founder-approved): /apps.swap2 4G, fstab-persisted, no restarts. Verification + closure note: https://projects.knownelement.com/issues/755#note-13
49 lines
3.8 KiB
Markdown
49 lines
3.8 KiB
Markdown
> Knowledge docs live on Discourse — perf write-up: [t/316](https://community.turnsys.com/t/316) ·
|
||
> [VP TechOps category](https://community.turnsys.com/c/vp-techops) ·
|
||
> Work tracking: [Redmine — Cloudron project](https://projects.knownelement.com/projects/88) ·
|
||
> Infra audit log: [t/298](https://community.turnsys.com/t/298)
|
||
|
||
# KNEL/cloudron — tsys-cloudron fleet-machine repo
|
||
|
||
Production Cloudron host **my.knownelement.com** (tailnet name
|
||
`tsys-cloudron.knel.net`). ~55 apps behind the Cloudron panel, platform
|
||
Postgres/MySQL/mail containers, monitored by Uptime Kuma (Cloudron group).
|
||
|
||
| Path | What |
|
||
|------|------|
|
||
| `perf/` | [#685](https://projects.knownelement.com/issues/685) perf work: memory census, pg role-cap SQL, applied sysctl, applied `daemon.json`, `apply-memory-shrink.sh` |
|
||
| `monitoring/` | Uptime Kuma Cloudron app coverage — 55-app inventory + sync checker (moved from PFVCluster, #435/#685/#727) |
|
||
| `ansible/` | Reserved — AWX/Ansible IaC convergence ([#454](https://projects.knownelement.com/issues/454)); playbooks land here as manual changes are codified |
|
||
| `scripts/` | Rule engine + git hooks (enforcement layer, ADOPTING pattern) |
|
||
| `tests/` | shellcheck wrapper (Docker — host stays clean) |
|
||
|
||
## Ground truth (2026-09-02, #685 — narrative + measurements on [t/316](https://community.turnsys.com/t/316))
|
||
|
||
- Platform Postgres slot exhaustion root-caused to 7 booking instances + Penpot
|
||
pinning ~61 idle connections each. Role caps applied (booking 10, Penpot 15,
|
||
every other app role 20), `max_connections` 500→700, `shared_buffers` 2GB,
|
||
idle-in-transaction timeout 15min. Victim `too many clients` FATALs → zero.
|
||
- Host sysctl: `vm.swappiness=10` (`perf/sysctl/`). THP already `never`, virtio
|
||
scheduler `none` — optimal, left alone.
|
||
- **`/etc/docker/daemon.json` APPLIED 2026-09-02** (#731): `live-restore: true`
|
||
(future dockerd restarts are zero-downtime) + `default-ulimits nofile
|
||
262144` (every container start gets real fd limits — the postgres prlimit
|
||
bridge is no longer needed; `perf/docker-daemon.json` is the SoR copy).
|
||
NOTE: this host's dockerd runs `--log-driver=journald` as a unit flag —
|
||
daemon.json must NOT set log-driver/log-opts or dockerd refuses to start
|
||
(bit us once, see #731). Container log rotation is journald's job here.
|
||
- **Memory shrink table APPLIED 2026-09-02** (#731) via the panel API
|
||
(`perf/apply-memory-shrink.sh`): 11 apps resliders, explicit limits
|
||
53.4 GB → 29.0 GB (~24 GB of the ~28 GB headroom reclaimed). Skips:
|
||
photos (manifest floor 3.5 GB), `bi.knownelement.com` + platform services
|
||
(do NOT shrink). All 159 apps verified running post-change. **2026-09-03 addendum:** OOM-driven bumps (thefnf 512M→1G, grafana 512M→1G, projects 600M→1.5G) + inotify-exhaustion dockerd/containerd incident — see #731 + [t/298](https://community.turnsys.com/t/298).
|
||
|
||
## Access
|
||
|
||
SSH `root@my.knownelement.com` — routed through the KNEL/PFVCluster
|
||
`tests/remote.sh` chokepoint (see AGENTS.md). Panel API token issued
|
||
2026-09-02 (`~/.creds/cloudron.env`, refs only in this repo) — prefer the
|
||
API over SSH for reads/limits; base image rebuilds (#710) still need SSH.
|
||
- **Inotify headroom APPLIED 2026-09-03** (#755, founder-approved): `fs.inotify.max_user_instances` 128→1024 + `max_user_watches` →1048576 in the same sysctl file, after containerd/dockerd died of inotify exhaustion (full fleet resync 09:13–09:51). Runtime-applied, no daemon restarts; error stream dry from 10:08. OOM fail set complete: thefnf/grafana/projects bumped (1G/1G/1.5G), zero OOMs post-recovery.
|
||
- **Swap 4G→8G APPLIED 2026-09-03** (#755, founder-approved): second swapfile `/apps.swap2` (4G, fstab-persisted, same-priority family) alongside the original `/apps.swap` — no swapoff, no restart, live `free -h` shows 8.0Gi. Idempotent applier: `perf/swap/grow-swap2.sh`. Headroom for the Monday app push.
|