DESIGN: budget + quota reality (flat-rate now, $200/mo after 09-22; flash-stretch routing)
This commit is contained in:
@@ -122,3 +122,90 @@ GLM fleet is wide; use it deliberately.
|
||||
with per-tier budget caps (LiteLLM vkeys already enforce spend).
|
||||
- Telemetry: every REPORT logs model used + tokens + cost via LiteLLM
|
||||
spend APIs, so routing decisions are auditable and tunable.
|
||||
|
||||
## Key proxy: placeholders only, material never leaves the vault (Charles, 2026-08-28 evening)
|
||||
- PROBLEM with existing gateways: they HOLD real keys in a webui/api and
|
||||
serve placeholders. We invert that: key material lives ONLY in Bitwarden
|
||||
(Secrets Manager) + HashiCorp Vault. A Go proxy (MOPAC/tools/keyproxy or
|
||||
bitwarden-go evolution) resolves placeholder -> real at the wire, holds
|
||||
material in memory only (fetch-on-demand, short TTL, never persisted,
|
||||
never in any admin UI; there is no admin surface — config only).
|
||||
- FRONTEND: consumers (harness, CLIs, containers) receive placeholder keys
|
||||
(mpk_... opaque refs). If one leaks, it's a placeholder: revoke/re-map in
|
||||
vault, zero upstream exposure. Defense in depth.
|
||||
- BACKEND connectors: Bitwarden Secrets Manager REST (machine accounts;
|
||||
AVOID the official SDK — source-available license conflicts with AGPL;
|
||||
plain REST via stdlib keeps AGPLv3 clean) + HashiCorp Vault KV v2 +
|
||||
AppRole (official Go api pkg is MPL-2.0, AGPL-compatible, vendor it).
|
||||
License: AGPLv3; MIT fallback only if a conflict still arises.
|
||||
- ADOPT crush's exposure-minimization protocol everywhere (per
|
||||
docs/PORTING-NOTES-crush.md + house practice): keys via os.environ/VAR
|
||||
indirection in configs (never literals), 0600 on secret files, redaction
|
||||
of key material in all logs/output, canonical-source regeneration
|
||||
(crush.json -> providers.json pattern), never exec env files inside
|
||||
scripting languages, session scrubbing.
|
||||
- Also feeds: study-secrets TASK (KNELSecretsManager notes) in flight.
|
||||
LiteLLM stays a sanctioned service exception (its z.ai key can later be
|
||||
injected at container start from the vault; not tonight).
|
||||
|
||||
## Tooling = standalone public FLOSS repos, loosely coupled (Charles, 2026-08-28 evening)
|
||||
- Dedicated CLIs/services each get their OWN public Gitea repo under the
|
||||
UKRRS org (org for ALL LLM work; AGPLv3), composable by anyone:
|
||||
ukrrs/mopac-keyproxy (created, seeded), later mopac-redmine-cli,
|
||||
mopac-discourse-cli, mopac-bitwarden-go, ... ukrrs/MOPAC (transferred
|
||||
2026-08-28) is the harness core. MOPAC references them; it does not
|
||||
contain them.
|
||||
- LOOSE COUPLING RULES (FLOSS without support burden): config-driven —
|
||||
no org-specific hosts/paths/defaults baked into code; contracts are
|
||||
exec + JSON in/out (Unix plumbing); each tool standalone-usable by a
|
||||
human; TSYS-specific policy/policy repos (agent-stack) live OUTSIDE the
|
||||
tools. We build for our needs first; genericness comes from clean
|
||||
contracts, not from supporting foreign requirements.
|
||||
|
||||
## ALL dev work in Docker (Charles, 2026-08-28 evening) — BIG RULE
|
||||
- No dev toolchains on the host: compiling, vetting, testing, hacking all
|
||||
happen inside containers (e.g. digest-pinned golang:1.26 builder with the
|
||||
repo bind-mounted, `go build/vet/test` inside, artifacts out via the
|
||||
mount). Host runs containers + screens; nothing else.
|
||||
- docker pull of anything needed is PRE-AUTHORIZED, no asking.
|
||||
- Exception cleanup: the harness-skeleton turn installed go1.26.7 to
|
||||
host ~/.local before this rule landed — remove it (follow-up TASK) once
|
||||
that turn completes; all future builds go through the builder container.
|
||||
- Rationale: reproducible, supply-chain-clean (CMMC posture), host stays
|
||||
pristine; same pattern as the LSP stack (digest-pinned containers).
|
||||
|
||||
## Events are V1 scope: webhook receiver now (Charles, 2026-08-28 evening)
|
||||
- NOT deferred. `mopac events` (Go HTTP receiver, in ukrrs/MOPAC) ships with
|
||||
the first build. The interaction layer (Redmine/Discourse/Gitea) must be
|
||||
able to PUNCH the harness: when Charles replies on Discourse, updates a
|
||||
Redmine ticket, or approves a Gitea PR, the harness knows and acts.
|
||||
- Trigger -> action semantics (dedup + replay, at-least-once):
|
||||
Redmine issue update/note on an active stack's issue -> dispatch/redirect
|
||||
turn for that stack; Discourse post reply in a stack's category/topic ->
|
||||
context update + response turn; Gitea PR approved/merged -> next pipeline
|
||||
step (build/test/deploy per repo config).
|
||||
- Signatures: Gitea HMAC-SHA256 secret header; Discourse webhook secret
|
||||
header; Redmine shared-secret header (verify, reject unsigned). Normalize
|
||||
to internal event record -> Redmine issue (SoR) -> conductor picks up.
|
||||
- Runs containerized (dev-in-docker rule), port published on host LAN;
|
||||
Cloudron apps get the webhook URL configured per-app.
|
||||
|
||||
## Hermes killed; OpenWebUI is the interactive front door (Charles, 2026-08-28 evening)
|
||||
- Hermes was Node — removed (dashboard/gateway/searxng, 2026-08-28 ~19:10).
|
||||
No Node middleware anywhere in MOPAC's path.
|
||||
- OWUI (Cloudron + SSO) becomes the ONLY interactive surface: MOPAC serves
|
||||
an OpenAI-compatible /v1/chat/completions endpoint (Go stdlib HTTP, in
|
||||
this repo, e.g. `mopac serve`). OWUI adds it as a connection; each stack/
|
||||
vertical registers as a "model". OWUI admin UI + Cloudron SSO = RBAC we
|
||||
do not build. Keys minted per OWUI connection (LiteLLM vkeys pattern).
|
||||
- Event path unchanged: Redmine/Discourse/Gitea webhooks -> `mopac events`
|
||||
(Go T1 receiver). OWUI chat and webhooks are two doors into the same loop.
|
||||
|
||||
## Budget + quota reality (Charles, 2026-08-28 evening)
|
||||
- Until 2026-09-22: legacy z.ai coding plan = FLAT RATE. Real budget = plan
|
||||
quota, not USD. Smaller/flash models stretch quota ~3x. Routing policy:
|
||||
grind/study/summarize -> flash tier (glm-5.3-flash now live on proxy,
|
||||
13 models); flagship (glm-5.3/5.2) only when the work needs it. Maximize
|
||||
usage of what we already pay for.
|
||||
- After 2026-09-22: org-wide AI cap $200/mo across all 8 stacks:
|
||||
2x $80 z.ai max plans + $40 OpenRouter. Design all budgets/vkeys to fit.
|
||||
|
||||
Reference in New Issue
Block a user