Verified after the 22:53 multiqueue bounce: NIC queues live at 4/4 and 2/2 (the whole point of the reboot), PVE config clean with nothing pending, both gateway lanes healthy, all 9 Open Terminal ports bound on the tailscale IP, PSI metrics fresh, and no removed relaunchers respawned. Moved the day-profile flip from 05:00 to 07:00 per human ruling (live timer re-armed + staged copy and docs synced; the gateway z.ai peak ladder is untouched). Also staged a ready-to-run script that drops each account's own Open Terminal key into its ~/.creds as openwebui.creds - not run yet; the master key list stays at ~/.creds/open-terminal.env. 💘 Generated with Crush Assisted-by: Crush:glm-5.2 [#610]
69 lines
3.9 KiB
Markdown
69 lines
3.9 KiB
Markdown
# NEXT — perf-opt closeout COMPLETE (2026-08-31); #610 open-terminal fleet DONE
|
||
|
||
## Tonight (2026-08-31 late, crush had the con)
|
||
- Agent-stack relaunchers REMOVED (6-): no screen/crush sessions auto-start
|
||
on reboot, any account. Backups: removed-agent-stacks/.
|
||
- Open Terminal fleet LIVE (#610): 9 accounts, tailscale-only
|
||
100.101.187.119:30000-30008, systemd open-terminal@<user>, keys in
|
||
/etc/ukrrs/open-terminal/ + ~/.creds/open-terminal.env. Ops note:
|
||
OPEN-TERMINAL.md. UAT PASSED (9-uat-openwebui.out).
|
||
- 22:00 night flip verified live (see TRACKING decisions).
|
||
- HUMAN OWED: wire production OpenWebUI per OPEN-TERMINAL.md. The
|
||
openwebui CONTAINER on this box (v0.3.10, :3000) is too old — upgrade to
|
||
0.11.x first. Cloudron prod (0.11.1) is fine.
|
||
- Multiqueue bounce DONE + VERIFIED (2026-09-01 morning audit): ens18
|
||
Combined 4/4, ens19 2/2; 4-host-netcheck.sh sweep clean (qm config
|
||
live, pending empty, bounce 22:51->22:53). Full stack healthy post-boot.
|
||
- 2026-09-01: night window now 22:00-07:00 (day-flip timer moved to 07:00).
|
||
|
||
## Final state — nothing owed (perf pass)
|
||
- All guest tuning live and verified (see 16:04-16:15 pass below).
|
||
- Host flags live: ssd=1/discard (ROTA=0), iothread, onboot=1, startup
|
||
order=10, cpuunits 9000/50, ide2 gone.
|
||
- Net multiqueue: queues=4/2 now IN live qm config (staged 16:40 by crush
|
||
via ssh); activates at the next VM start — ONE reboot still owed, human-
|
||
timed: run ~/projects/ultix/5-activate-multiqueue.sh (kills crush
|
||
sessions, ~4 min). HUMAN 2026-08-31 eve: nothing else is running on this
|
||
system — bounce THIS SESSION, right after the remaining session tasks.
|
||
Guest oneshot ukrrs-net-multiqueue auto-runs ethtool -L on that boot.
|
||
- Gateway boot-race + live-restore endpoint loss: permanently fixed by
|
||
ukrrs-gateway-ensure unit (enabled, both lanes).
|
||
- Root cause of the 15:59 no-op: the script copy executed then predates the
|
||
queues= fix (silent same-value rewrite); current config verified correct.
|
||
|
||
## Verified live after the 16:01 bounce (all read-only from the guest)
|
||
- Disks: ROTA=0, discard 4K/1G on sda+sdb+sdc → ssd=1/discard=on ACTIVE ✅
|
||
- Kernel: bbr+fq, 16M socket buffers, dirty_bytes 1G/256M, min_free 384M,
|
||
aio 1M, inotify 512/1M, port range 10240-65535, slow_start_after_idle=0,
|
||
tcp_tw_reuse=1, THP=madvise, DefaultLimitNOFILE 65536 ✅
|
||
- Mounts /,/data1,/data2 = noatime; swap sda5 prio -2 unused ✅
|
||
- Units: all ukrrs slices + day/night + psi + builder-prune + thp units
|
||
installed and enabled; PSI textfile writing (16:06) ✅
|
||
- Docker: daemon.json fully live (live-restore, log caps 20m×3, 172.16/12
|
||
pools, metrics :9323); 21 containers up ✅
|
||
- Desktop: sddm autologin worked (Relogin=true), session alive on seat0/tty2,
|
||
greeter burn GONE ✅
|
||
- Gateway prod was DOWN at boot (bind race + live-restore endpoint loss) →
|
||
recovered via compose --force-recreate; /status = mode normal, serving ✅
|
||
|
||
## Nothing owed. Optional sanity check any time (crush can ssh now):
|
||
~/projects/ultix/4-host-netcheck.sh
|
||
|
||
## Noted, no action taken
|
||
- mopac-demo/mcli fake containers: restart=no, exited at the first bounce,
|
||
left down (test stubs; your call).
|
||
- Prometheus/harness daemons not deployed here yet (that is OPT-9's wiring).
|
||
|
||
## HARDWARE SLIPPED TO ~OCT 2026 (ruling 2026-08-31)
|
||
- PRESUME 8 vCPU / 48G on this VM until the #601 window: interim fences
|
||
(accounts 3G/4.5G) and the daynight interim values ARE the operating
|
||
config; the commented post-upgrade block stays dormant.
|
||
- Gateway cpuset hard reservation (16-17) waits for the window → the
|
||
compose cgroup_parent wiring is THE starvation fix meanwhile (ticketed).
|
||
- q35 / 20 vCPU / 128G / GPUs / USB3 reseat / NFS-outage-for--02 all move
|
||
to the same October window (#601 → version october-2026-window, due
|
||
10-31; #606).
|
||
- Fri 2026-09-04 = #609: troubleshooting-only outage (USB SDR reseat +
|
||
zero-cost socket/PSU/BIOS checks). Host down ⇒ this VM down, auto-starts
|
||
via onboot=1; gateway auto-recovers (validated 2026-08-31).
|