Buy the labor, not the judgment
By Keith Townsend · July 22, 2026
The pitch is that a local bug-fix agent needs a frontier tier in the loop to escalate to. I rebuilt the chain as originally intended: one loop, one deterministic test gate, one feedback contract, run unchanged across 70 real bug-fix pull requests, swapping only the model in the worker and controller seats and metering every call. A frontier model diagnosing the local worker's failures, as the loop's controller, recovered no net task a free deterministic feedback loop did not. The same frontier model doing the labor, as the worker, cleared work the local model could not. So the intelligence belongs in the worker seat, and the test still decides done.
Measured on deterministic coding repair, where a test harness gives an unfalsifiable pass/fail, and after the benchmark supplied file-level localization. The rulings are about where the determines-done authority and the required capability tier sit in an agentic loop. The open edges are autonomous repository debugging, domains without an executable evaluator, and larger replication.
Do run the local model as the worker and expect it to clear the solvable share. Most first-attempt misses were execution failures a stateful loop with deterministic feedback recovered for free, not gaps a bigger model was needed to close.
Do make the deterministic evaluator the authority. The test harness decides accept or reject, drives the loop, and supplies the feedback. No model judges its own work or another's.
Do when the hard tail needs a higher tier, spend it as the worker, not the controller. In one harness a frontier model doing the labor took recovery from 14 to 20; the same model merely diagnosing added nothing.
Don’t buy the frontier as cheap control over a weak local worker. A worker that cannot write the patch cannot act on the diagnosis. That quadrant paid the most and recovered the least.
Don’t read a single tuned loop as a general control result. This one held across 70 heterogeneous defects without bespoke orchestration. That is the claim, not one demonstration.
Self-funded editorial. No vendor paid for this answer, reviewed it, or saw it before publication. The free local worker is Gemma, a Google model, run on owned NVIDIA hardware; the hosted arms were metered OpenAI calls, and an adjacent run, held out of the comparison, used a flat Anthropic subscription. Disclosure: Google Cloud is a client of The CTO Advisor LLC; NVIDIA, OpenAI, and Anthropic are not. The ruling is the author's alone.
Does a local bug-fix agent need a frontier model in the loop to escalate to? That's the pitch, and it sounds responsible. Keep a cheap local worker for the routine repairs, and buy a frontier model's judgment as the loop's controller: it reads each failure, diagnoses what went wrong, and steers the small model home. You get frontier quality at local prices. I rebuilt the chain as originally intended to test exactly that, and the answer inverted on me. The frontier model earned nothing as the controller. It earned its keep as the worker. The seat is the finding.
The bench was deliberately boring. Seventy real bug-fix pull requests, mined by git archaeology, each shipping the fix's own test suite as an unfalsifiable pass/fail gate. One stateful repair loop, one deterministic validator, one feedback contract, held constant across all of it. Four arms shared that single harness: a model emits a search-and-replace edit, the deterministic gate replies with the test result, and the loop iterates up to four times. Only the model in the worker and controller seats changed, and every hosted call was metered. The free worker was a local 26B reasoning model, Gemma, running on a DGX I own. The hosted arms swapped GPT-5 mini and GPT-5.5 into the two seats. One honest caveat up front: the benchmark supplies the touched files, so this measures localized repair, not repository search. And one discipline that makes the numbers mean anything: the loop was never rewritten around a defect. It held across 70 heterogeneous bugs without bespoke orchestration, and that is the claim, not one demonstration.
The free loop did the heavy lifting
Before any paid model enters the picture, the local worker with rich deterministic feedback recovered 14 of the 22 tasks it missed on the first attempt. Read that again. Most first-attempt misses weren't capability gaps a bigger brain needed to close. They were execution failures a stateful loop recovered for free, on nothing but the test harness telling the model precisely what broke. This builds on the original loop-control lab rather than replacing it. That lab proved things this one never tests: a deterministic model at temperature zero retries a repair loop to the identical wrong answer, and reviewing every output makes accuracy worse. It also identified the regime this lab lives in, a reasoning model that can produce a different attempt given the same feedback. Both labs stand. The question here is the next one of that regime: where does the frontier tier earn its place?
The controller seat bought nothing
Not in the control seat. A GPT-5.5 controller reading each failure and advising the same local worker recovered no net task the free loop did not. That held across three seeds, 13-13, 10-10, and 11-10 on the tasks both arms ran cleanly, never once ahead, for about ninety-eight thousand frontier tokens a run spent to draw even. I gave the controller a fairer shot too. The per-turn version saw only the current failure, so a variant got the full trajectory of attempts and prior guidance. It nudged to 13-11 on one 14-task subset, a faint edge inside the sampling band. That's the only sign in the whole run that a controller might add anything, and it's a lead, not a finding. One more seed decides whether it's real.
Here's the obvious objection: isn't this just "use a bigger model"? Of course a frontier worker solved more. Yes, as the worker, and that's the inversion. The pitch was to avoid the expensive thing, frontier labor, by buying frontier judgment instead. That arrangement bought nothing. Moving the same frontier model out of the control seat and into the worker seat, same harness, took recovery from 14 to 20 of 22. The lever was never the loop or the diagnosis. It was which model held the pen. A worker that can't write the patch can't act on the diagnosis, which is why the frontier-controller-over-weak-worker quadrant paid the most and recovered the least. If the hard tail needs a higher tier, spend it as the worker. And through all four arms, the deterministic evaluator kept the only authority that matters: the test harness decides accept or reject, drives the loop, and supplies the feedback. No model judges its own work or another's.
How you buy tokens picks the seat
The economics refused to stay a footnote. Metered, the frontier is something you reserve for the hard tail, and the meter has teeth: the metered account hit its quota and stopped the experiment mid-run. An adjacent arm on a flat Anthropic subscription ran through. On a flat plan the frontier's labor is free up to the cap, so you run it as the worker and the whole reserve-it-for-the-tail logic dissolves. The pricing regime is a first-class variable in where you place the tier, not an afterthought for the finance team.
That adjacent arm deserves its own honesty. A fifth run cleared all 22 tasks, but it swapped my constrained harness for Claude Code, an agentic loop with tools and file exploration. It changed two variables at once, the harness and the model, so it proves neither, and I held it out of the comparison. I nearly conflated it with the model-tier ladder before a second look caught it. The clean experiment was to hold Claude Code fixed and drop in a less capable model, and that lab has since run. Lab 012 re-ran the fifth arm metered with per-task usage capture, 22 of 22 again at $13.79 across 10.96 million tokens, and then the separation came back against the tier: a mini model in the same agentic harness matched all 22 for $3.14, after plateauing at 16 and 17 in this lab's constrained loop. The harness was the lever. That result belongs to Lab 012; what this lab could claim on its own was only that the conflation was a trap.
What this lab did not prove is a shorter and important list. Three seeds of no controller advantage is not an equivalence proof, and the trajectory-aware controller's faint edge is a single subset inside the noise. The benchmark handed over file-level localization, 21 of 22 tasks a single file, though those files ran a median of about 2,600 lines. So nothing here rules out a frontier tier helping where fault localization is in play, or in domains with no executable evaluator at all. Autonomous repository debugging is an open edge, not a settled one.
So the verdict I'd put in front of an Enterprise Architect is the title. Buy the labor, not the judgment. Run the local model as the worker and expect the free deterministic loop to clear the solvable share. When the residual tail justifies a frontier tier, put that tier in the worker seat, where it can hold the pen, and let how you buy tokens decide how often it sits there. What you should not buy is a frontier model's diagnosis of a worker that can't execute it. And whichever model sits in whichever seat, done is not the model's call. The test decides. The next time a vendor pitches you an escalation tier for your agent loop, ask the only question this bench turned out to be about: which seat does it sit in, and who decides done?
The numbers
Where each layer belongs
| Layer | Placement |
|---|---|
Layer 0 · Compute Compute & Network Fabric The local first-attempt tier. A free 26B reasoning worker runs here on a DGX and clears the solvable share on its own, which is why the paid tiers buy less than the pitch claims. | Retained |
Layer 2B · Runtime Application Runtime & Execution The test harness and the loop orchestrator are deterministic code you run, not a model. This gate decides accept or reject, drives the loop, and supplies the feedback. | Retained |
Layer 2C · Reasoning Agentic Infrastructure — The Reasoning Plane The determines-done authority stays with the deterministic evaluator, never the model. Repair capability is delegated to the worker, and the finding is that the worker's tier, not a frontier controller diagnosing it, is what clears the hard tail. Buy the frontier's labor here, not its judgment. | Retained |
Assessments at the time of the lab
Method and disclosure
Self-funded editorial. The harness is loopcontrolbench, published open under the MIT license at github.com/kltownsend/loopcontrolbench: the reproduce-or-drop gate, the sandboxed runner, and the arm definitions. The vendor assessments and the layer placements this evidence feeds stay proprietary.
The essay on this page was rendered from the lab’s frozen structured record by a model writing under this practice’s voice specification, with a deterministic gate checking every figure against the record before publication. The record is the canonical surface: where prose and record disagree, the record rules.
Download the raw lab detail (Markdown)