Buy the labor, not the judgment
The pitch is that a local bug-fix agent needs a frontier tier in the loop to escalate to. I rebuilt the chain as originally intended: one loop, one deterministic test gate, one feedback contract, run unchanged across 70 real bug-fix pull requests, swapping only the model in the worker and controller seats and metering every call. A frontier model diagnosing the local worker's failures, as the loop's controller, recovered no net task a free deterministic feedback loop did not. The same frontier model doing the labor, as the worker, cleared work the local model could not. So the intelligence belongs in the worker seat, and the test still decides done.
By Keith Townsend · July 22, 2026
The verdict
Measured on deterministic coding repair, where a test harness gives an unfalsifiable pass/fail, and after the benchmark supplied file-level localization. The rulings are about where the determines-done authority and the required capability tier sit in an agentic loop. The open edges are autonomous repository debugging, domains without an executable evaluator, and larger replication.
Self-funded editorial. No vendor paid for this answer, reviewed it, or saw it before publication. The free local worker is Gemma, a Google model, run on owned NVIDIA hardware; the hosted arms were metered OpenAI calls, and an adjacent run, held out of the comparison, used a flat Anthropic subscription. Disclosure: Google Cloud is a client of The CTO Advisor LLC; NVIDIA, OpenAI, and Anthropic are not. The ruling is the author's alone.
Video
How I know
The pitch under test: a local bug-fix agent needs a frontier tier in the loop to escalate to. The re-run held one loop architecture, one deterministic test gate, and one feedback contract constant across 70 real bug-fix pull requests, and swapped only the model in the worker and controller seats.
The free local worker with stateful deterministic feedback recovered 14 of the 22 tasks it first missed. Adding a frontier model as the controller diagnosing those failures recovered no net task the free loop did not, 13 to 13 on the tasks both arms ran cleanly. The controller was the wrong place to spend intelligence.
The right place was the worker seat. In the same harness, moving from the local 26B to a frontier worker took recovery from 14 to 20 and cleared most of a hard residual tail. A weak worker cannot act on a good diagnosis. The tier has to hold the pen.
The economics split by how you buy tokens. Metered, reserve the frontier for the tail. On a flat subscription its labor is free up to the cap, so run it as the worker. The metered account hit its quota and stopped the experiment mid-run; the subscription arm did not.
What the bench measured
The detail
Four arms shared one harness: one model emits a search-and-replace edit, a deterministic gate replies with the test result, and the loop iterates up to four times. Only the model in the worker and controller seats changed. That shared harness is what makes the four comparable.
The free local worker recovered 14 of its 22 first-attempt misses on rich deterministic feedback alone. A GPT-5.5 controller reading each failure and advising the same worker recovered no additional matched task, and it held across three seeds: 13-13, 10-10, and 11-10 on the tasks both arms ran cleanly, never once ahead, for about ninety-eight thousand frontier tokens a run spent to draw even. That controller saw only the current failure. Handed the full trajectory instead, a fair variant, it nudged to 13-11 on one subset, a faint edge inside the sampling band, worth another seed but not yet a result.
Moving the frontier model out of the control seat and into the worker seat, same harness, took recovery to 20 of 22 and cleared three of the five hardest residual tasks. The lever was never the loop or the diagnosis. It was which model held the pen. A weak worker cannot execute a good diagnosis, so buying the frontier's judgment as control is buying the wrong thing.
A fifth run cleared all 22, but it swapped the harness for Claude Code, an agentic loop with tools and file exploration, so it is held out of the comparison. It changed two variables at once and proves neither. Whether the agentic harness or the model did that work is a separate lab, and the clean way to run it is to hold Claude Code fixed and drop in a less capable model. That lab has since run. Lab 012 re-ran this fifth arm on the metered API with per-task usage capture: 22 of 22 again, $13.79, 10.96 million tokens. And the separation came back against the tier: a mini model in the same harness matched all 22 for $3.14. The harness was the lever.
Is this not just "use a bigger model"? Of course a frontier worker solved more.
Yes, as the worker, and that is the inversion. The pitch was to keep a cheap local worker and buy the frontier's judgment as loop control. That is the arrangement that bought nothing, 13 to 13 on the tasks both arms ran cleanly. The frontier only helped when it did the labor itself, which is the expensive thing the pitch was trying to avoid.
And the cost flips with how you buy tokens. Metered, you reserve the frontier for the hard tail. On a flat subscription its labor is free up to the cap, so you run it as the worker. During these runs the metered account hit its quota and stopped the experiment mid-run; the flat-plan arm ran through. The pricing regime is a first-class variable, not a footnote.
Where each layer belongs
| Layer | Placement |
|---|---|
Layer 0 · Compute Compute & Network Fabric The local first-attempt tier. A free 26B reasoning worker runs here on a DGX and clears the solvable share on its own, which is why the paid tiers buy less than the pitch claims. | Retained |
Layer 2B · Runtime Application Runtime & Execution The test harness and the loop orchestrator are deterministic code you run, not a model. This gate decides accept or reject, drives the loop, and supplies the feedback. | Retained |
Layer 2C · Reasoning Agentic Infrastructure — The Reasoning Plane The determines-done authority stays with the deterministic evaluator, never the model. Repair capability is delegated to the worker, and the finding is that the worker's tier, not a frontier controller diagnosing it, is what clears the hard tail. Buy the frontier's labor here, not its judgment. | Retained |
The two questions this lab now knows to ask
The in-harness ladder located a graduated tail: a hosted mini worker under frontier control cleared one of five residual tasks, a frontier worker three. Within this harness the last two held out. How the boundary moves when localization is withheld or the task class hardens is open.
An adjacent run put a top model inside Claude Code, a different agentic harness with tools and its own loop, and it cleared the whole set including the last two. But that swapped the harness as well as the model, so it proves neither. The next lab holds Claude Code fixed and drives it with a less capable model. If a weaker model in the same agentic harness still clears the tail, the harness was the lever; if recovery collapses toward these numbers, the model tier was carrying it. Answered: Lab 012 ran exactly that, and the harness was the lever. A mini-tier model inside Claude Code cleared all 22, twice, while plateauing at 16 and 17 in this lab's constrained loop. The Opus arm, re-run metered with real usage capture, came in at $13.79 across 10.96 million tokens; the mini finished identical work for $3.14.
The per-turn controller tied the free loop across three seeds. Handed the full history of attempts and prior guidance instead, it edged the free loop 13 to 11 on one matched subset. That is the only sign in the whole run that a controller might add anything, and it is a single trajectory inside the noise band. One more seed decides whether it is real.
What it did not prove
- The per-turn controller result holds across three seeds, 13-13, 10-10, and 11-10 on matched tasks, so it is no advantage seen in three runs, not one. That is still not an equivalence proof. The trajectory-aware controller's faint edge, 13-11 on a single 14-task subset, sits inside the sampling band and is a lead, not a finding.
- The benchmark supplied file-level localization, 21 of 22 tasks a single file, though those files ran a median of about 2,600 lines. This measures localized repair, not autonomous repository debugging, and cannot rule out that a frontier helps where fault localization is in play.
- The Claude Code result is adjacent, not a rung. It changed the harness and the model at once, so it attributes to neither, and any read of it as clean model separation is wrong. Lab 012 has since made the separation cleanly, and this caution stands as the record of what this lab could and could not claim on its own.
Notes from the author, Keith Townsend
This builds on the original loop-control lab, it does not replace it. That one proved things this one never tests: that a deterministic model at temperature zero retries a repair loop to the identical wrong answer, and that reviewing every output makes accuracy worse. It also identified the regime this lab lives in, a reasoning model that can produce a different attempt given the same feedback. What I ask here is the next question of that regime, where the frontier model earns its place, as the loop's controller or as its worker. Both labs stand.
The result I did not expect was how much the framing turned on where the frontier model sat. As a controller it earned nothing. As the worker it earned its keep. I nearly conflated a fifth run, a top model inside Claude Code, with the model-tier ladder, and it took a second look to see that swapping the harness is not swapping the model. That confusion is the next lab, not this one.
Assessments at the time of the lab
Method and disclosure
Self-funded editorial. The harness is loopcontrolbench, published open under the MIT license at github.com/kltownsend/loopcontrolbench: the reproduce-or-drop gate, the sandboxed runner, and the arm definitions. The vendor assessments and the layer placements this evidence feeds stay proprietary.