# ultix-streaming optimization โ€” tracking HUD One table, updated in place at each checkpoint. States: done โœ… / doing ๐Ÿ”„ / blocked โ›” / next โณ | id | lane | state | item | |---|---|---|---| | OPT-1 | guest | โœ… | recon: system + ukrrs fleet (REPORT ยง1) | | OPT-2 | guest | โœ… | REPORT.md + staged configs + questions-v1.md | | OPT-3 | host | โœ… | audit parsed โ†’ REPORT ยง8 (T7500, CPU2 empty, clean IOMMU, Q4000 skip) | | OPT-4 | host | โณ | #601 ~OCT 2026: 2ร— X5675 + 6ร— 16GB โ†’ 24T/192G (NFS outage for -02 first; USB3 reseat; q35/20c/128G). PRESUME 8 vCPU/48G until then | | OPT-5 | host | โœ… | one-shot landed + verified post-reboot; net multiqueue in qm config, activates at next VM start (bounce approved for tonight) | | OPT-6 | guest | โœ… | guest prep applied + verified (bbr, THP, slices, timers, noatime, daemon.json, autologin) | | OPT-7 | guest | โณ | #603: root growth MANUAL ONLY, GROW-ROOT-RUNBOOK.md, user-executed | | OPT-8 | repos | โœ… | PMO backpressure design note in ukrrs/docs/harness; code = ticketed | | OPT-9 | integ | โณ | #604: node_exporter svc + gateway pressure scrape (Q10) | | OPT-10 | integ | โณ | #605: k8s join (Q11, blocked by #601) + proxmox-ctl token (Q12); GPUs #606 | | #610 | fleet | โœ… | Open Terminal 9-account fleet (tailscale-only :30000-30008) + OpenWebUI UAT PASS | Inbox (mid-task interrupts): none. Decisions log (latest wins): - 2026-08-31 (night, con handed to crush): #610 DONE in one session: open-terminal 0.11.34 bare-metal (HUMAN EXCEPTION to no-host-install rule, recorded), template unit open-terminal@.service, per-user TOML keys (never in ps/repo; ~/.creds holds copies), bind TAILSCALE-ONLY 100.101.187.119 per human ruling (0.0.0.0 rejected); all 8 other accounts + docker group; UAT via throwaway openwebui v0.11.1 (matches Cloudron prod) ALL PASS incl. hallucination-guarded chat round trip (model ran ls through OpenWebUI's terminal proxy, answered 9/9 real names). NOTE: prod openwebui container (v0.3.10, :3000) is too old for Open Terminal โ€” upgrade to 0.11.x before wiring; Cloudron 0.11.1 is fine. Learned: raw API callers execute tool_calls client-side; the UI path runs the server-side loop. - 2026-08-31 (night): 22:00 night-profile flip OBSERVED LIVE from the guest (dirty_bytes 1Gโ†’2G, bg 256Mโ†’512M, /data2 readahead 128Kโ†’1M; day timer restores at 05:00). First runtime-only flip since install. - 2026-08-31 (night): agent-stack relaunchers REMOVED by 6-remove-agent-stacks.sh (reachableceo-agent-stack.service, tsg-supervisor.timer/.service 5-min self-heal, dead tsg-agent-stacks.service; launchers + units backed up to removed-agent-stacks/). No screen/crush sessions auto-start on reboot anymore; human: stacks no longer used for pmo/work. - 2026-08-31 (night): docs synced post-pass: REPORT v1.1 (status banner, per-item STATUS tags, ยง6/ยง8 truth), RUNBOOK amended for the Oct replan, all paths ~/optimize โ†’ ~/projects/ultix (repo moved). Human ruling: system idle tonight โ†’ multiqueue bounce approved for this session, after remaining tasks. Agent-stack teardown prepped as 6-remove-agent-stacks.sh (kills boot-time RCEO/TSG screen respawn + the 5-min supervisor self-heal; unit/launcher backups to removed-agent-stacks/; root also quits leftover root screen ukrrs-gateway-final, leaves 'boot' alone). - 2026-08-31 (late): Redmine replan complete: #601 Oct window (due 10-31, High, version october-2026-window; Redmine API refuses start_date on update โ€” Oct-1 start recorded in the note instead), NEW #608 cgroup_parent wiring (High; THE starvation fix on 8 vCPU), NEW #609 Fri 09-04 troubleshooting-only outage (USB SDR reseat + socket/PSU/BIOS checks; NO upgrades; host down โ†’ this VM auto-starts via onboot=1, gateway auto-recovers). #605/#606 noted. docs repo pushed (main 9eb324e). - 2026-08-31 (eve): HARDWARE WINDOW โ†’ ~OCT 2026; presume 8 vCPU/48G operating values; cpuset reservation deferred โ†’ cgroup_parent wiring = the starvation fix; one VM-5111 reboot owed for multiqueue (5-activate-multiqueue.sh, human-timed; crush may roll OTHER VMs on pfv-tsys5 anytime). - 2026-08-31 (eve): PUSH AS YOU GO standing permission (human monitors gitea/redmine/discourse, not ssh/crush). House rules updated + pushed (86d8565). - 2026-08-31 16:5x: Redmine = record: perf pass + RCA filed as #602 (project 55, tracker Support, left New for human review); open items #601 (Fri window, Urgent, due 09-04) #603 root growth #604 metrics #605 k8s+proxmox (blocked by #601) #606 GPUs (relates #601). mred = docker exec wrapper (~/.local/bin/mred โ†’ mopac-cli-mred container). - 2026-08-31 16:40: net multiqueue CLOSED: queues=4/2 written to live qm config via crush ssh (user authorized non-disruptive host mods); activates at next VM start; guest oneshot enables channels on that boot. Earlier silent no-op explained: the 15:59 run used a pre-fix script copy. - 2026-08-31 16:3x: crush CAN ssh to the host from this workspace (old "hard-banned" note obsolete); still human-gated for disruptive ops. - 2026-08-31 (post-reboot validation): prod gateway down at boot ROOT-CAUSED twice over: (a) port bind to tailscale IP lost the dockerd-vs-tailscaled race (docker never retries failed starts), (b) live-restore dropped the container's network endpoint (host resolver โ†’ gateway-db unresolvable โ†’ crash loop). Recovered via compose --force-recreate (healthy, mode=normal). Permanent fix staged: ukrrs-gateway-ensure unit, BOTH lanes, MUST be enabled before Friday's bounces. - 2026-08-31: multiqueue host flags verified live (ROTA=0 + discard 1G on all 3 disks), but guest still runs 1 queue/NIC: ethtool not installed; 3-post-reboot-fixes.sh staged (installs ethtool + oneshot unit, sets 4/2 queues). - 2026-08-31: mopac-demo/mcli fake containers left DOWN (restart=no, exited at first bounce); prometheus/harness daemons not deployed on this box (pre-existing; belongs to OPT-9). - 2026-08-31: STORAGE OPS MANUAL ONLY: boot-time auto-grow service removed, qm resize removed from the one-shot, grow-root.sh deleted. Runbook = GROW-ROOT-RUNBOOK.md, every command typed by the human. - 2026-08-31: KDE autologin into UNLOCKED always-alive session (Relogin=true); autolock OFF (instant Jump/iPad attach outranks lock). XFCE only if Sunshine capture disappoints post-GPU. - 2026-08-31: streamlined to 1-guest-prep.sh + 2-host-one-shot.sh (one reboot wave; no post-reboot steps; no storage). - 2026-08-31: constraint: host pfv-tsys5 NEVER reboots casually (NFS server for -02); Friday CPU swap needs a planned NFS outage first. - 2026-08-31: GPUs are COMPUTE cards; Quadro 4000 skipped (Fermi: no NVENC, no kernel-6 driver, no modern CUDA); clean IOMMU groups verified. - 2026-08-31: day/night = runtime-only cgroup flips 22:00/05:00; NO cpufreq driver exists host or guest. - 2026-08-31: zram rejected; no docker data-root move; skip tuned (sysctl.d + units). - 2026-08-31: second USB3 SDR card missing on host (only one Renesas visible); Friday reseat/check. ----