ops(archive): TASK prompts and REPORT verifications, studies through V1

This commit is contained in:
2026-08-29 05:22:43 -05:00
parent bd0cea1757
commit b2c4b74d76
37 changed files with 2049 additions and 0 deletions
+126
View File
@@ -0,0 +1,126 @@
# REPORT — `harness serve`: OpenAI-compatible front door (OWUI is the client)
Date: 2026-08-29 ~01:30 | Repo: ukrrs/MOPAC (harness/) | Status: DONE, live on LAN port 8090, pushed to origin main (2901bb8..704f905)
## What it is
`harness serve` = the OWUI front door from DESIGN (hermes killed; OpenWebUI =
interactive surface with Cloudron SSO + admin RBAC we do not build). OWUI
adds MOPAC as an OpenAI connection; each task class registers as a "model".
OWUI chat and webhooks (`harness events`, port 4100) are now two doors into
the same conductor loop. v0: stateless, non-streaming, tools off.
## Endpoints (port 8090, own port — coexists with events on 4100)
| Route | Auth | Behavior |
|---|---|---|
| `POST /v1/chat/completions` | Bearer vkey | ONE bounded stateless conductor turn over the full history OWUI sends; reply = single assistant message + usage. `temperature`/`max_tokens` forwarded. `stream:true` = explicit 400. |
| `GET /v1/models` | Bearer vkey | the servable catalog (below) |
| `GET /healthz` | none | liveness |
Auth = constant-time SHA-256 digest compare of the Bearer vkey; missing /
wrong / malformed keys all get one byte-identical generic 401. vkey is a
config ref (`[serve] vkey_ref = "env:HARNESS_SERVE_VKEY"`), never logged.
Unknown model = 400 naming the valid ones. Upstream failures = terse 502.
## Model catalog (as configured in harness.toml right now)
Every `[models.classes]` class is exposed as model `mopac-<class>`, routed
through the existing `[models]` tier table to the concrete proxy model:
| OWUI model name | Tier | Concrete model (sent to LiteLLM) |
|---|---|---|
| `mopac-primary` | mopac-primary | glm-5.3 (default / flagship+) |
| `mopac-code` | mopac-code | glm-5.2 (flagship) |
| `mopac-architecture` | mopac-code | glm-5.2 |
| `mopac-review` | mopac-review | glm-5-turbo (mid) |
| `mopac-summarize` | mopac-review | glm-5-turbo |
| `mopac-writeback` | mopac-review | glm-5-turbo |
| `mopac-study` | mopac-study | glm-4.7-flash (flash) |
| `mopac-read` | mopac-study | glm-4.7-flash |
| `mopac-vision` | mopac-vision | glm-4.6v |
`[serve] enabled_models` can narrow this to a subset if you want fewer
buttons in OWUI (validated at config load).
## OWUI connection settings (plug these in)
- **Base URL**: `http://192.168.3.78:8090/v1`
- **API key (vkey)**: `mopac-owui-f2fa1879c71b8d7c83098c7838f8859d3b85d3bf`
(canonical copy: `~/.coordinate/secrets/mopac-serve-vkey.env`, 0600 —
same value as env var `HARNESS_SERVE_VKEY`; revoke = mint a new one +
restart)
- In OWUI: Admin Settings -> Connections -> add OpenAI API connection with
the above, then pull the model list (the 9 `mopac-*` models appear);
per-user/group model access is OWUI's RBAC, not ours.
- OWUI must NOT request streaming for this connection (v0 returns an
explicit 400 on `stream:true`; OWUI tolerates non-streaming providers —
toggle in the connection settings if a client defaults to streaming).
## How to run / restart
```sh
cd ~/projects/meta/MOPAC/harness
set -a; . ~/.coordinate/secrets/mopac-harness-vkey.env # HARNESS_LITELLM_KEY
. ~/.coordinate/secrets/mopac-serve-vkey.env # HARNESS_SERVE_VKEY
set +a
./dev.sh serve # container mopac-serve, port 8090 on the LAN
# stop: docker rm -f mopac-serve (exact container, never pkill)
```
It is RUNNING now (container `mopac-serve`, up as of this report).
## Implementation notes
- Turn machinery REUSED, not copied: `loop.turn` refactored into a shared
bounded-turn core (`runTurn`) that takes arbitrary message history, an
optional tool palette, and forwarded max_tokens/temperature. `once`/`loop`
behavior unchanged (all prior tests pass untouched); the serve path calls
the same core with tools hard-off (a hallucinated tool call is refused as
a tool result, never executed).
- History assembly: OWUI's own leading system message is preserved verbatim
(OWUI personas win); only when none is sent does the harness prepend a
minimal vertical-identity prompt.
- Commits (logical chunks, all Docker-built via digest-pinned builder):
8614827 loop refactor, c54a5a4 [serve] config, c9e86ee serve + CLI +
dev.sh runner, 704f905 docs. Pushed to origin main.
## Test results
`./dev.sh check` (build + vet + test, golang:1.26 builder): ALL GREEN —
```
ok internal/config ok internal/events ok internal/intake
ok internal/llm ok internal/loop ok internal/models
ok internal/serve ok internal/tools ok internal/writeback
```
- Serve tests: scripted fake OpenAI upstream against the REAL server — auth
matrix (missing/malformed/wrong key 401, generic body, no vkey leak),
catalog + enabled_models subset, end-to-end chat (concrete model on the
wire, no tools, OpenAI response shape, usage), history assembly both
ways, multi-round usage accounting (refused tool call fed back, sums
40/30/70), temperature/max_tokens forwarding, unknown-model 400 naming
valid ones, stream 400, empty/bad-role/malformed 400s, upstream 502,
fail-fast construction.
- Loop tests: ServeTurn system-prompt prepend/preserve, tools-off refusal
(allow-listed `pwd` NOT executed), knob forwarding, Router accessor.
- Live LAN proof (python urllib, curl banned on host): healthz 200; wrong
key 401; catalog 200 (all 9 models); unknown model 400 listing them;
stream 400; REAL turns through LiteLLM — `mopac-primary` -> glm-5.3
("MOPAC demo door" reply, 255 tokens, finish stop) and `mopac-study` ->
glm-4.7-flash (166 tokens). Audit lines log routing + counters only.
## Next (per TASK: streaming, tools; plus the obvious follow-ons)
1. **Streaming**: SSE `stream:true` path on the serve door (needs the llm
client's streaming + truncation-resume work, already on the phase 3
list).
2. **Tools in serve turns**: per-vertical gated tool palette inside serve
turns (the machinery already supports a palette; it is deliberately nil
for v0).
3. Per-connection vkeys (LiteLLM-vkey-style minting via keyproxy) instead
of one vkey per harness instance.
4. Each onstage/offstage stack gets its own `harness serve` instance with
its own vertical name + class map — that is how "each stack registers
as a model in OWUI" scales past one box.