# REPORT — `harness serve`: OpenAI-compatible front door (OWUI is the client) Date: 2026-08-29 ~01:30 | Repo: ukrrs/MOPAC (harness/) | Status: DONE, live on LAN port 8090, pushed to origin main (2901bb8..704f905) ## What it is `harness serve` = the OWUI front door from DESIGN (hermes killed; OpenWebUI = interactive surface with Cloudron SSO + admin RBAC we do not build). OWUI adds MOPAC as an OpenAI connection; each task class registers as a "model". OWUI chat and webhooks (`harness events`, port 4100) are now two doors into the same conductor loop. v0: stateless, non-streaming, tools off. ## Endpoints (port 8090, own port — coexists with events on 4100) | Route | Auth | Behavior | |---|---|---| | `POST /v1/chat/completions` | Bearer vkey | ONE bounded stateless conductor turn over the full history OWUI sends; reply = single assistant message + usage. `temperature`/`max_tokens` forwarded. `stream:true` = explicit 400. | | `GET /v1/models` | Bearer vkey | the servable catalog (below) | | `GET /healthz` | none | liveness | Auth = constant-time SHA-256 digest compare of the Bearer vkey; missing / wrong / malformed keys all get one byte-identical generic 401. vkey is a config ref (`[serve] vkey_ref = "env:HARNESS_SERVE_VKEY"`), never logged. Unknown model = 400 naming the valid ones. Upstream failures = terse 502. ## Model catalog (as configured in harness.toml right now) Every `[models.classes]` class is exposed as model `mopac-`, routed through the existing `[models]` tier table to the concrete proxy model: | OWUI model name | Tier | Concrete model (sent to LiteLLM) | |---|---|---| | `mopac-primary` | mopac-primary | glm-5.3 (default / flagship+) | | `mopac-code` | mopac-code | glm-5.2 (flagship) | | `mopac-architecture` | mopac-code | glm-5.2 | | `mopac-review` | mopac-review | glm-5-turbo (mid) | | `mopac-summarize` | mopac-review | glm-5-turbo | | `mopac-writeback` | mopac-review | glm-5-turbo | | `mopac-study` | mopac-study | glm-4.7-flash (flash) | | `mopac-read` | mopac-study | glm-4.7-flash | | `mopac-vision` | mopac-vision | glm-4.6v | `[serve] enabled_models` can narrow this to a subset if you want fewer buttons in OWUI (validated at config load). ## OWUI connection settings (plug these in) - **Base URL**: `http://192.168.3.78:8090/v1` - **API key (vkey)**: `mopac-owui-f2fa1879c71b8d7c83098c7838f8859d3b85d3bf` (canonical copy: `~/.coordinate/secrets/mopac-serve-vkey.env`, 0600 — same value as env var `HARNESS_SERVE_VKEY`; revoke = mint a new one + restart) - In OWUI: Admin Settings -> Connections -> add OpenAI API connection with the above, then pull the model list (the 9 `mopac-*` models appear); per-user/group model access is OWUI's RBAC, not ours. - OWUI must NOT request streaming for this connection (v0 returns an explicit 400 on `stream:true`; OWUI tolerates non-streaming providers — toggle in the connection settings if a client defaults to streaming). ## How to run / restart ```sh cd ~/projects/meta/MOPAC/harness set -a; . ~/.coordinate/secrets/mopac-harness-vkey.env # HARNESS_LITELLM_KEY . ~/.coordinate/secrets/mopac-serve-vkey.env # HARNESS_SERVE_VKEY set +a ./dev.sh serve # container mopac-serve, port 8090 on the LAN # stop: docker rm -f mopac-serve (exact container, never pkill) ``` It is RUNNING now (container `mopac-serve`, up as of this report). ## Implementation notes - Turn machinery REUSED, not copied: `loop.turn` refactored into a shared bounded-turn core (`runTurn`) that takes arbitrary message history, an optional tool palette, and forwarded max_tokens/temperature. `once`/`loop` behavior unchanged (all prior tests pass untouched); the serve path calls the same core with tools hard-off (a hallucinated tool call is refused as a tool result, never executed). - History assembly: OWUI's own leading system message is preserved verbatim (OWUI personas win); only when none is sent does the harness prepend a minimal vertical-identity prompt. - Commits (logical chunks, all Docker-built via digest-pinned builder): 8614827 loop refactor, c54a5a4 [serve] config, c9e86ee serve + CLI + dev.sh runner, 704f905 docs. Pushed to origin main. ## Test results `./dev.sh check` (build + vet + test, golang:1.26 builder): ALL GREEN — ``` ok internal/config ok internal/events ok internal/intake ok internal/llm ok internal/loop ok internal/models ok internal/serve ok internal/tools ok internal/writeback ``` - Serve tests: scripted fake OpenAI upstream against the REAL server — auth matrix (missing/malformed/wrong key 401, generic body, no vkey leak), catalog + enabled_models subset, end-to-end chat (concrete model on the wire, no tools, OpenAI response shape, usage), history assembly both ways, multi-round usage accounting (refused tool call fed back, sums 40/30/70), temperature/max_tokens forwarding, unknown-model 400 naming valid ones, stream 400, empty/bad-role/malformed 400s, upstream 502, fail-fast construction. - Loop tests: ServeTurn system-prompt prepend/preserve, tools-off refusal (allow-listed `pwd` NOT executed), knob forwarding, Router accessor. - Live LAN proof (python urllib, curl banned on host): healthz 200; wrong key 401; catalog 200 (all 9 models); unknown model 400 listing them; stream 400; REAL turns through LiteLLM — `mopac-primary` -> glm-5.3 ("MOPAC demo door" reply, 255 tokens, finish stop) and `mopac-study` -> glm-4.7-flash (166 tokens). Audit lines log routing + counters only. ## Next (per TASK: streaming, tools; plus the obvious follow-ons) 1. **Streaming**: SSE `stream:true` path on the serve door (needs the llm client's streaming + truncation-resume work, already on the phase 3 list). 2. **Tools in serve turns**: per-vertical gated tool palette inside serve turns (the machinery already supports a palette; it is deliberately nil for v0). 3. Per-connection vkeys (LiteLLM-vkey-style minting via keyproxy) instead of one vkey per harness instance. 4. Each onstage/offstage stack gets its own `harness serve` instance with its own vertical name + class map — that is how "each stack registers as a model in OWUI" scales past one box.