From b69d3872039f610eab3bbd946cf20d8462c2aa30 Mon Sep 17 00:00:00 2001 From: reachableceo Date: Fri, 28 Aug 2026 18:58:49 -0500 Subject: [PATCH] DESIGN: MVP demo bar + model selection layer (tier map, LiteLLM aliases, auto-pick v1) --- DESIGN.md | 29 +++++++++++++++++++++++++++++ 1 file changed, 29 insertions(+) diff --git a/DESIGN.md b/DESIGN.md index f18577f..5aef0b9 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -93,3 +93,32 @@ by humans/legacy stacks — interaction layer must not require the harness; interface can come later without redesign. MCP servers we run get replaced by these Go CLIs behind a minimal MCP adapter, killing the docker/python MCP fleet. + +## MVP demo bar (Charles, 2026-08-28 ~19:00) +Fire up the harness (screen or docker) and run a turn whose prompt is +"tell me about yourself" routed through LiteLLM; the GLM 5.3 self- +description comes back and lands as the REPORT. That is the MVP. The +0800 checklist (compile, dry-run, real single-turn, LSP wired, studies) +builds toward it; the demo itself is this one exchange. + +## Model selection layer (Charles, 2026-08-28 ~19:00) +Claude-Code-style model picking for the harness and its subagents: +the harness chooses the right model per body of work (overall + per-task). +GLM fleet is wide; use it deliberately. +- Inventory on proxy today (12): glm-4.5{,-air,v}, 4.6{,v}, 4.7{,-flash}, + 5, 5-turbo, 5.1, 5.2, 5.3. No glm-5.3-flash configured yet; flash-class + today = glm-4.7-flash. Adding 5.3-flash later = one LiteLLM model_list + entry, zero code change (see aliases). +- v0 (tonight, skeleton): static TASK CLASS -> MODEL TIER map in + harness.toml. Classes: study/read -> flash tier; code/architecture -> + flagship (glm-5.2/5.3); review/summarize/writeback -> mid (5-turbo/5.1); + vision when needed (4.6v/4.5v). Config-only, no heuristic code. +- Aliases are the contract: harness requests `mopac-study`, `mopac-code`, + `mopac-review` ... LiteLLM resolves alias -> concrete model. Swapping + models (e.g. 4.7-flash -> 5.3-flash when live) = proxy config edit only. +- v1 (post-pilot): auto-selection — cheap flash model classifies each + incoming TASK (Redmine issue text) and picks tier + model before the + main turn spends anything; subagent spawns inherit class-based routing + with per-tier budget caps (LiteLLM vkeys already enforce spend). +- Telemetry: every REPORT logs model used + tokens + cost via LiteLLM + spend APIs, so routing decisions are auditable and tunable.