One expensive model decides and tastes. Cheaper models do the cooking, each at its own station. Tests check every plate before anyone tastes it, and I approve anything that leaves the building.
Part 5 of the AI Workflow series build log
I pick the work and approve anything that leaves the building. Nothing emails, posts or runs on my Windows machine without my OK.
Claude can write Gmail drafts but can't send them (send is blocked in its settings). OpenClaw asks me before every Windows command; Muse asks before every email it sends.
One Claude Code session on the expensive model keeps the plan and the results, never the working details. A long session is the single biggest cost, so every task starts fresh.
Claude Code, Opus. Each turn re-reads the whole conversation: one long day cost ~$49, most of it context. Habit: a new session per task, /clear between jobs, never let the root grow fat.
For a batch of small, well-defined changes, the head chef hands the whole batch to the line. Each cook works in its own copy of the repo, so nobody bumps elbows.
director workflow: Plan (Opus, read-only) → one git worktree per task → Haiku or Sonnet worker → repo's own tests in Docker (a Haiku helper runs them; the pass/fail is parsed by code) → fail: retry, then a stronger model → blind Opus review, at most 2 revise rounds → merge + push. First real run: 5 tasks, 4 right first time, ~41% cheaper than doing it all on Opus.
We tested the taster too: a wrong measurement that still passed every test got slipped in on purpose. The reviewer caught it.
A model that only answers yes or no, very fast and very cheap. It's good at narrow questions and useless at broad ones, so it only watches for now.
TypeSafe's Jev, ~$0.0001 per check. Asked "is 5/8 plywood written as `actual: 19.0 / 32`?": 0.99 on the right code, 0.03 on a planted bug. Asked "does this meet every criterion?": 0.72 vs 0.70. Shadow mode: its answer is logged beside the Opus verdict. Across two runs it agreed with Opus 6 of 6. It sees only privacy-gated, redacted text, and it is skipped for client repos. If they keep agreeing, it can skip reviews.
Think of a tasting panel of volunteers who work for free. Seven of them taste the same plate at once, and a doorman decides first what each one is allowed to see. For now they only watch, beside the Opus taster.
Experts, all free, run in parallel: Groq gpt-oss-120b · Groq Qwen 3.8 27B · Cloudflare Workers AI Qwen2.5-Coder 32B Gemini 3 Flash (public only) · Antigravity CLI (public only) OpenRouter free Nemotron Ultra (public only) a local Qwen2.5-Coder 7B on Ollama, on the GPU (any repo) About 20 to 45 s per review. Privacy gate: secrets block the run, the rest is redacted first. By repo class: public = all experts; own private = only experts that don't train on inputs (Groq, Cloudflare, local); client = the local model only. It earns its way in: 15+ items with zero false passes and at least 90% agreement. After that it does first-round reviews, and Opus reviews only when it objects. First evidence: Gemini wrongly rejected a correct change twice; Groq and the local model got it right.
When the head chef needs to know something, it sends a researcher and gets back a short summary, not the pile of files they read.
Search and exploration helpers run on Sonnet (the built-in explorer is overridden to Sonnet, with a check that warns when the built-in changes upstream).
My personal assistant runs my calendar and mail. It and Claude leave each other notes in my Gmail, and I can read every one.
Claude → Muse: drafts titled [CC>MUSE], read in Muse's 7:45 am briefing. Muse → Claude: emails to myself titled [MUSE>CC], trusted only if Gmail marks them Sent from my account. Notes are information, never orders; no client details go in them.
A tray app on my PC that can pop a notification or run a short command, like reading what hardware I have. It's boxed in and asks me every time.
Four allowed actions (notify, run, info, approvals) out of 52 the tray offers. Commands run sandboxed, no network, 30 s limit. Its own agent now runs on a free Gemini Flash model. Jev is the decision model inside OpenClaw, but it hasn't been exercised there yet. Its master key is off-limits to Claude.
The older way I work is a crew of terminal windows, one per role. Each role now runs the model that fits it, and a morning crew manager shows its plan before it starts anything.
tmux assembly line: 26 roles + an Operator. A model per role: 13 Opus, 12 Sonnet, 1 Haiku. Morning crew manager: plan first, then --apply. Planned: clear a window as its memory fills; later a Jev check that reads each window (done / stuck / waiting) so the Operator only wakes for real decisions. MoltBot, an older agent that co-hosts on Twitch, joins over the local network (planned).
Every new session reads a short index of lessons and rules first, so fresh sessions start smart instead of starting over.
A memory index read at every start, a to-do list, lessons written back when something goes wrong, and a backup of the setup to a public repo that never holds keys or client files.
A receipt book kept per client: how much each client's work used, and whether the hours were mine or the agents'.
Tokens per client, each reply counted once. "My time" comes from my own prompts, set against agent time. Also counted: commits, and the free panel's tokens. It updates at each session start. Transcripts are kept 90 days; monthly summaries go to a private hub repo.