On 3 October I asked one Claude to audit the rest. For about two days I’d been pushing OldMate, ManyCP and TrustAudience through the same loop: Opus plans, Sonnet builds, a fresh Opus reviews, round again. It had run 108 times and I wanted to know what it was costing me.
The builder was the cheap part. The expensive parts were the planning step, the review rounds, and a few stalls that had nothing to do with any model. Two of the biggest costs were defaults I’d set myself.
This is as of October 2026, on Opus 5.5 and Sonnet 5.5, and model names date fast. The numbers are agent-tallied from a small sample, so read this as a diary with receipts, not a study.
What is the Opus plan, Sonnet implement loop?
Each agent is a Claude Code subagent: a markdown file with a model and an effort level in its frontmatter, started with a fresh context every time. A workflow script wires them together.
- The lead (Opus 5.5, high effort) checks the brief, answers the builder’s questions, and triages the reviewer’s findings: fix, decline, out of scope or ask me.
- The implementer (Sonnet 5.5, high by default; 15 of 49 builds ran at medium) builds from the brief, runs the checks, and stops to ask rather than guess. I dropped it from xhigh on 1 October on Anthropic’s effort guidance, not an A/B of my own.
- The reviewer (Opus 5.5) is a new one every round. It sees the goal, the criteria, the diff and the checks, never the conversation.
- The organiser: a Claude thread I talk to, which turns requests into briefs, launches the loop and reports back.
It ends when nothing above minor is left, or after six rounds. It began with a plan unless the brief switched that off, and 16 of the runs that built did.
Where do the tokens go in an AI coding agent loop?
All 108 runs, by step, from the run records:
| Step | Share of tokens | Share of agent time |
|---|---|---|
| First build (Sonnet) | 20% | 30% |
| Plan (Opus lead) | 16% | 15% |
| Review and triage (Opus) | 34% | 30% |
| Fixes the review asked for (Sonnet) | 21% | 21% |
| Final checks (Sonnet) | 7% | 3% |
| Everything else | 2% | 1% |
That’s 550 agent calls, about 57 million tokens and 67 agent-hours. Review plus the fixes it asked for took 55% of the tokens and half the time. The first build, the bit everyone talks about, took a fifth of the tokens.
Antoine Buteau’s write-up of the January Tokenomics paper found the same shape: code review used 59.4% of tokens and coding 8.6% (30 ChatDev tasks, GPT-5), so compare shapes, not digits. I haven’t converted tokens to dollars, and I don’t know how the records count cached input.
46 of my 108 runs stopped at the plan step
If the lead had a question it thought was mine, the run stopped and handed it back. That happened 46 times: 94 questions, 5.3 million tokens and 309 agent-minutes, about 116,000 tokens and seven minutes each.
Every question came with a default. Of the 43 relaunches, 31 started within two minutes, so the default got taken, usually by the organiser thread rather than me. And 42 of the 43 planned again, so more than half of everything spent on planning (5.3 of 9.3 million tokens) bought a plan that was almost always redone.
I had the audit agent sort the 94 questions one by one. About 35% were genuinely mine: a brief’s US$3 budget couldn’t pay for the eval it asked for. About 30% were details the lead should have decided, like which of Apple’s three logo sizes to use. About 20% were scope fences the brief had already settled. About 8% were real bugs in the brief: a grep that had to come back empty while two banned strings sat in OpenAI’s own docs. That’s what planning is for, but Sonnet can catch that kind of disagreement later, and cheaper. It stopped to ask 12 times across all 108 runs, and the lead’s answers cost 0.32 million tokens in total.
Does an Opus plan make Sonnet’s first build better?
Is planning even useful, when models mostly work this stuff out as they go? The runs that built, split by whether they planned:
| Runs that built | First review found a blocker or major | Median run | |
|---|---|---|---|
| Planned | 34 | 17 of 32 (53%) | 791k tokens, 49 min |
| No plan | 16 | 6 of 14 (43%) | 324k tokens, 18 min |
This isn’t a clean experiment. The unplanned runs were mostly follow-ups: 10 of the 16 were on one OldMate sign-in feature (fixes, polish, rebases and two four-hour redesigns). What I can say is narrower. I found no sign the plan made Sonnet’s first draft better. Planned runs did take about 2.4 times the tokens and 2.7 times the minutes, but most of that gap is that they were bigger jobs: the plan step itself was a median 114k tokens.
So after the audit I made planning opt-in. Sonnet builds straight from the brief, and a plan happens only if the brief asks for one: work that spans several apps, touches money, privacy or auth, or lands in code nobody has mapped. Only money, privacy or legal claims, credentials or security, and anything irreversible or public can stop a run to ask me. The lead settles everything else with a default and lists it for me to veto at ship time. The handful of runs since is too few to judge quality.
Same builder, three repos, very different first reviews
46 runs built something and got a first review, and 23 had a blocker or major the lead confirmed after reading the code: OldMate 17 of 20, ManyCP 3 of 9, TrustAudience 3 of 17.
I’d like to say the codebase matters more than the model. My numbers mostly say the reviewer does. 15 of OldMate’s 20 reviews came from a stricter design reviewer (Opus at higher effort, with the app’s design rules), and all 15 found a blocker or major. With the standard reviewer OldMate was 2 of 5, about where ManyCP sat. I can’t separate a harder codebase from a pickier reviewer, but the builder stayed put and the number moved. TrustAudience, a small Cloudflare Worker app with no AGENTS.md at the time, had the cleanest first builds, so rule files don’t explain it either.
What did Sonnet get wrong? Nothing subtle: a view where the first three acceptance criteria weren’t built, a type-check that looked green only because turbo replayed a cached result, about 80 stale proof screenshots. All reported done anyway, a cousin of the tests-pass-but-the-code-is-broken bugs I collected here. So the builder now reports each done-when item as met or not, with evidence, and anything unmet goes back before an Opus reviewer is spent on it. That’s from 3 October, so I have no numbers on it yet.
The most expensive review was mine
On 2 October I looked at the screenshots of OldMate’s new sign-in and sign-up screens and typed “this looks horribly cluttered.” They had just come out of one fix loop (93 minutes) and were partway through another (197). The redesign that followed dropped the stepper, the offer card and the consent tick above the provider buttons. The Apple and Google button artwork survived and the layout didn’t, so call it about five loop-hours of fixes on a screen I then replaced, before the redesign runs themselves (272 and 258 minutes).
No plan, test or reviewer would have caught it. A mockup I’d approved first would have. So a new look for a screen now needs a reference I’ve approved, named in the brief, before the build starts. If there isn’t one, the first run makes one for my yes.
Seven review rounds on a small audio fix
A few days earlier, before the loop was a script, I’d run the same pattern by hand with no limit. In a DimeTown session on 28 September I told the agent to keep looping independent reviews and fixes until only minor nits remained. One workstream was a console full of this:
RangeError: Failed to execute 'setValueAtTime' on 'AudioParam': Time must be a finite non-negative number: -0.2
A mood-change whoosh was booked 0.3 seconds before the downbeat, but on a brand-new audio context the first bar is only 0.1 seconds away, so the start time came out at minus 0.2. The first commit was 19 lines added and 4 removed in one file, plus a 128-line test. Round 1’s reviewer found nothing above minor. By my own rule that was the end. The loop ran six more rounds over the next three hours.
Two of those rounds found a blocker or a major, and both came from the previous round’s fixes. The fixes for round 2 broke 8 of 58 screenshot tests, and a test changed in round 6 would still pass if you trimmed the list it was checking. The commit that shipped touched four files, 705 insertions and 20 deletions, 597 of them tests. At 16:30 I asked “what has caused all this churn?”
The agent’s answer started with “Mostly my process.” Across that session, each fresh reviewer re-read everything with no memory of what was already settled, so each dug one layer deeper, and every minor sent the loop round again. The fix had been briefed as “audio starts on every browser”, which pulled in iOS unlocking, taps versus swipes and modifier keys. Its fix: stop after two rounds, keep fixes to the reported bug, merge before stacking, agree pass and fail numbers up front.
The script I wrote three days later takes the spirit but not the number. It caps at six rounds, stops after any round that fixes only minors, and refuses a done-when item like “polish the settings screen” unless it names a test, a number or a screenshot. Six is a ceiling: of the 41 runs that reached ready, 32 got there in two rounds or fewer.
Delays that weren’t the model
- The permission classifier denied about 40 actions across the three repos between 29 September and 3 October, and approving in chat doesn’t lift one. A nine-minute change to a deploy workflow took 14 hours 35 minutes from the denial at 22:43 to the push the next afternoon.
- The Mac slept for about an hour on 1 October and every agent in flight stalled. The thread that launches loops now asks the app to keep it awake.
If you want to build this loop
Three agent files in ~/.claude/agents/. These are the lines that matter:
# loop-lead.md
model: opus
effort: high
tools: Read, Glob, Grep, Bash # no Edit or Write: it reads and decides
# loop-implementer.md
model: sonnet
effort: high
# loop-reviewer.md
model: opus
effort: high
tools: Read, Glob, Grep, Bash
A dynamic workflow script calls them in order (the lead plans only if asked, the implementer builds, a new reviewer reviews, the lead triages, the implementer fixes) and owns the round counter, so no agent decides when to stop.
A brief has a one-sentence outcome, two to five done-when items a reviewer or test can answer yes or no, an out-of-scope item, decisions with defaults, the checks to run and the proof to capture. The rules I’d copy:
- Stop only for money, privacy or legal claims, credentials or security, or anything irreversible or public.
- Plan only when the brief asks.
- Cap the rounds at six, and stop after any round that fixes only minors.
- Refuse vague words (polish, robust, clean up, nicer, seamless, better, properly) unless something checkable sits beside them.
- The builder reports each done-when item as met or not, with evidence.
- Verify with build caches bypassed, a conflict-marker scan, and proof newer than the last edit.
- Get a mockup approved before any new look.
Methods, and how far to trust this
- The runs. 108 feature-loop runs that finished between the morning of 1 October and midday on 3 October 2026, Brisbane time: OldMate 48, TrustAudience 38, ManyCP 22. Counts, tokens and minutes come from the run records Claude Code writes for each workflow, tallied by script and re-run before publishing. A run isn’t a feature (OldMate’s sign-in took 23), and “blocker or major” means one the lead confirmed after reading the code.
- The rest. An agent sorted the 94 questions, so treat the split as rough. The DimeTown numbers come from the session transcript and the fix’s git history. The sample is about two days, one developer, my own repos.