Files
ultix/NEXT.md
T
mrcharles c810942239 Open Terminal fleet for all accounts + OpenWebUI UAT pass (#610)
open-terminal 0.11.34 bare-metal (human-authorized host install) with a
systemd template unit: reachableceo on tailscale-only :30000 (cwd
~/projects) and all 8 other accounts on :30001-30008, each with its own
key kept out of ps in per-user TOMLs, all users added to the docker group.
Keys live in ~/.creds + /etc/ukrrs/open-terminal, never in this repo.

UAT: throwaway OpenWebUI v0.11.1 (matched to the human's Cloudron prod)
driven purely over its HTTP API against the beta gateway lane - connection
verify/config PASS, direct exec as uid 1001 PASS, and the chat round trip
PASS: the model emitted run_command, we executed it through OpenWebUI's
terminal proxy (server holds the key), and the final answer named 9/9 real
project directories; earlier negative runs prove the test catches
hallucinated output. Ops note OPEN-TERMINAL.md documents production wiring
(prod openwebui container v0.3.10 is too old and needs an upgrade first).

Also tonight: agent-stack relaunchers removed (6-, backups kept) so no
screen/crush sessions auto-start on reboot; 22:00 night-profile flip
observed live; docs synced for the earlier teardown.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
[#610]
2026-08-31 22:50:53 -05:00

3.9 KiB
Raw Blame History

NEXT — perf-opt closeout COMPLETE (2026-08-31); #610 open-terminal fleet DONE

Tonight (2026-08-31 late, crush had the con)

  • Agent-stack relaunchers REMOVED (6-): no screen/crush sessions auto-start on reboot, any account. Backups: removed-agent-stacks/.
  • Open Terminal fleet LIVE (#610): 9 accounts, tailscale-only 100.101.187.119:30000-30008, systemd open-terminal@, keys in /etc/ukrrs/open-terminal/ + ~/.creds/open-terminal.env. Ops note: OPEN-TERMINAL.md. UAT PASSED (9-uat-openwebui.out).
  • 22:00 night flip verified live (see TRACKING decisions).
  • HUMAN OWED: wire production OpenWebUI per OPEN-TERMINAL.md. The openwebui CONTAINER on this box (v0.3.10, :3000) is too old — upgrade to 0.11.x first. Cloudron prod (0.11.1) is fine.
  • LAST ACTION of the session: multiqueue bounce via ~/projects/ultix/5-activate-multiqueue.sh (approved; kills crush sessions ~4 min). After reboot verify: ethtool -l ens18 -> Combined 4/4; ens19 -> 2/2; then ~/projects/ultix/4-host-netcheck.sh for the full sweep.

Final state — nothing owed (perf pass)

  • All guest tuning live and verified (see 16:04-16:15 pass below).
  • Host flags live: ssd=1/discard (ROTA=0), iothread, onboot=1, startup order=10, cpuunits 9000/50, ide2 gone.
  • Net multiqueue: queues=4/2 now IN live qm config (staged 16:40 by crush via ssh); activates at the next VM start — ONE reboot still owed, human- timed: run ~/projects/ultix/5-activate-multiqueue.sh (kills crush sessions, ~4 min). HUMAN 2026-08-31 eve: nothing else is running on this system — bounce THIS SESSION, right after the remaining session tasks. Guest oneshot ukrrs-net-multiqueue auto-runs ethtool -L on that boot.
  • Gateway boot-race + live-restore endpoint loss: permanently fixed by ukrrs-gateway-ensure unit (enabled, both lanes).
  • Root cause of the 15:59 no-op: the script copy executed then predates the queues= fix (silent same-value rewrite); current config verified correct.

Verified live after the 16:01 bounce (all read-only from the guest)

  • Disks: ROTA=0, discard 4K/1G on sda+sdb+sdc → ssd=1/discard=on ACTIVE
  • Kernel: bbr+fq, 16M socket buffers, dirty_bytes 1G/256M, min_free 384M, aio 1M, inotify 512/1M, port range 10240-65535, slow_start_after_idle=0, tcp_tw_reuse=1, THP=madvise, DefaultLimitNOFILE 65536
  • Mounts /,/data1,/data2 = noatime; swap sda5 prio -2 unused
  • Units: all ukrrs slices + day/night + psi + builder-prune + thp units installed and enabled; PSI textfile writing (16:06)
  • Docker: daemon.json fully live (live-restore, log caps 20m×3, 172.16/12 pools, metrics :9323); 21 containers up
  • Desktop: sddm autologin worked (Relogin=true), session alive on seat0/tty2, greeter burn GONE
  • Gateway prod was DOWN at boot (bind race + live-restore endpoint loss) → recovered via compose --force-recreate; /status = mode normal, serving

Nothing owed. Optional sanity check any time (crush can ssh now):

~/projects/ultix/4-host-netcheck.sh

Noted, no action taken

  • mopac-demo/mcli fake containers: restart=no, exited at the first bounce, left down (test stubs; your call).
  • Prometheus/harness daemons not deployed here yet (that is OPT-9's wiring).

HARDWARE SLIPPED TO ~OCT 2026 (ruling 2026-08-31)

  • PRESUME 8 vCPU / 48G on this VM until the #601 window: interim fences (accounts 3G/4.5G) and the daynight interim values ARE the operating config; the commented post-upgrade block stays dormant.
  • Gateway cpuset hard reservation (16-17) waits for the window → the compose cgroup_parent wiring is THE starvation fix meanwhile (ticketed).
  • q35 / 20 vCPU / 128G / GPUs / USB3 reseat / NFS-outage-for--02 all move to the same October window (#601 → version october-2026-window, due 10-31; #606).
  • Fri 2026-09-04 = #609: troubleshooting-only outage (USB SDR reseat + zero-cost socket/PSU/BIOS checks). Host down ⇒ this VM down, auto-starts via onboot=1; gateway auto-recovers (validated 2026-08-31).