Lab 012 · Editorial lab

Buy the harness, not the tier

An architect escalating repair work has rungs available: a free local loop, a paid mid-tier model in that same loop, then a frontier model inside an agentic harness. I priced all of them. Same 22 certified bug-fix repairs, same deterministic test gate, six models from a local Gemma to two generations of Opus, run both as a constrained loop and inside Claude Code. The free local loop clears 17. The paid mid-tier loop clears 16 and 17 across two runs, no better. The harness is what closes the rest, and it does that for a mini model as readily as a frontier one. So which rung is actually load-bearing, and what does each one cost?

By Keith Townsend · July 26, 2026

The call

The verdict

Measured on localized repair with an executable test, one task pool of 22, one week of vendor pricing. The rulings are about where capability enters an agentic stack and what each rung of an escalation ladder costs per verified unit of output. The open edges are repository-scale debugging without localization, domains without an executable evaluator, and variance bounds beyond two runs per arm.

Dotriage on capacity you already carry. The free local loop clears three quarters of this work at zero marginal cost, and idle capacity comes in three forms: owned hardware, unused cloud-spend commitment, and unused subscription headroom.
Doescalate what the gate hands back into an agentic harness, and put the cheapest instrument-literate model in it. The mini cleared all 22 for $3.14; two Opus generations cleared the same 22 for four times that.
Dorun the ladder in both directions. Escalate when the clock binds. Route overflow back to idle hardware when the budget binds. The gate makes both directions safe because a verified repair is verified regardless of which rung produced it.
Don’tbuy a better model for the constrained loop. The paid mid-tier loop scored 16 and 17 across two runs against the free local model's 17 and 17. Same apparatus, no separation from free.
Don’trent GPU to self-host the escalation rung. Agentic sessions are VRAM-bound, not compute-bound, so the rented class costs more per verified repair than the API and delivers less than half the capability. You cannot rent your way to efficiency on this class of model.
Don’tread token counts as a business metric. Three models spanning a fourfold price range landed within 13 percent of each other on tokens, and one model run twice against itself varied 24 percent. Cost per verified repair is the number that decides anything.

Self-funded editorial. No vendor paid for this answer, reviewed it, or saw it before publication. The free local worker is Gemma, a Google model, run on owned NVIDIA hardware; the paid arms were metered OpenAI and Anthropic calls, and the rented-GPU arms ran on AWS. Disclosure: Google Cloud is a client of The CTO Advisor LLC; NVIDIA, OpenAI, Anthropic, and AWS are not. The ruling is the author's alone.

The walkthrough

Video

Video slot — sponsored labs fill this with a series.
The bench

How I know

The pitch under test: routing repair work to a cheaper model saves money, and the counter-pitch that cheap models flail and cost more by the end. Both failed. Across 14 arms on the same 22 gate-verified repairs, capability entered through the apparatus, not the price column, and the cost of finished work tracked the rate card alone.

The ladder, priced per verified repair: free local triage clears 17 of 22 at roughly zero. A paid mid-tier model in the same loop clears 16 and 17 across two runs, no separation from free. The same mini inside Claude Code clears all 22 for $3.14. Two Opus generations clear the identical 22 for $13.79 and $12.56. Escalating only the five the gate names costs $1.12.

Renting GPU to self-host the escalation rung fails on both axes at once. Agentic sessions park context in VRAM while barely touching compute, so concurrency is bought with memory, and the box already on the bench holds most of a rented fleet's sessions for nothing. The rented arm cost more than the API and solved fewer than half the tasks.

The transferable rule is a dispatcher, not a ladder. Spend committed capacity first, at every rung it exists: idle hardware, unused cloud commitment, unused subscription headroom. Escalate to metered capability only for what the gate hands back, buy the minimum tier that clears the bar, and route overflow back down when the budget binds instead of the clock.

What the bench measured

Configurations that cleared all 22
gpt-5.4-mini twice, Opus 4.8, Opus 5, all inside the same agentic harness
4
Cost per verified repair, mini harness vs Opus
$3.14 against $13.79 for identical finished work; both directly measured
$0.14 vs $0.63
Free local loop, two runs
paid mid-tier in the same loop: 16 and 17, no separation from free
17 and 17 of 22
Token spread across a 4x price range
inside the mini's own 24% run-to-run variance; the harness sets the token budget, the tier multiplies the rate
13%
Rented-GPU self-host of the escalation rung
at $4-6 of machine time: more than the API, for under half the capability
9 of 22
Total spend, every meter, mistakes included
$73.75 rented infrastructure, $47 and change in metered tokens, including $5.45 lost to a mid-run credit exhaustion and ~$24 of idle GPU; a reproducer pays for their failures too
~$121

The detail

Start on capacity you already carry. A Gemma 4 31B in a constrained loop clears 17 of the 22 on deterministic test feedback alone, and the gate hands back the exact five it could not solve. The earlier Xeon lab priced that rung: about nine cents per verified repair on rented Granite Rapids, six at the measured Savings Plan rate, and approaching zero on owned idle hardware or inside a commitment you already bought and did not use. This is not a story about owning a Spark. Unused cloud-spend commitment is idle capacity too.

The next rung up is the one that doesn't work. A paid mid-tier model in that same constrained loop scored 16 of 22, and 17 on a second run, against the free local model's 17 and 17. Two runs each, dead parity. Buying a better model without changing the apparatus around it bought no separation from free. That is the result that reframes the ladder, because it means the money is not in the tier.

It's in the harness. The same mini model that plateaued at 16 and 17 in the loop clears all 22 inside Claude Code. Six tasks recovered from apparatus, not parameters. That is where capability actually enters, and it is not exotic: read the repository, run the tests, read the failure, try again, with tools instead of a fixed protocol. Escalating just the residual five through it costs $1.12 against $12.56 for pushing all 22 through a frontier model, an eleven-fold spread for the same finished work. And even that $1.12 is a spot price. A flat-rate subscription seat with weekly headroom left over is idle capacity at this rung, the same way the Xeon lab's spare cycles were idle capacity at the one below. Idle Claude tokens can do real work too, and the marginal cost of the residual five inside an already-paid seat rounds to zero until the cap binds.

So should you buy the GPU class you can actually get? I priced it, and no. Agentic sessions are bought with memory, not compute: each one parks its context for its whole life while barely touching the die, so the honest unit is dollars per concurrent session per hour. Four entry instances give about nineteen sessions for $7.44 an hour. The box already on my bench holds fifteen for nothing. Renting the liquid class buys four more sessions at seven dollars an hour, which is not a purchase, it is a rounding error with a meter attached. The tier above is worse on the same axis and quota-locked besides, and it sells memory bandwidth against a per-task latency that is serial and irreducible. If you do rent, scale out with the smallest instance rather than up: consolidated multi-GPU boxes charge you for interconnect that independent sessions never touch.

What separates the rungs is what a retry costs, and that decides which model belongs on each. On idle capacity a retry is free, so a variance-prone cheap model is not a liability, it is a method: let it grind and let the gate certify each attempt. The Xeon lab audited 585 scored attempts and found no validator pass that failed its held-out checks, which is what makes grinding safe rather than reckless. On a metered API the arithmetic inverts. Every retry bills, so what you are actually buying up the ladder is first-pass success. That is the same property as instrument literacy, seen from the invoice.

One constraint keeps you honest, and it is why the escalation rung has to be bought rather than substituted. You cannot simply wrap the free local model in the harness to skip the paid step: Gemma 4 31B goes from 17 down to 11 inside Claude Code. The harness multiplies a model that can operate its instruments and taxes one that cannot. Instrument literacy is its own capability and it does not track parameter count. Above that bar the tier is close to irrelevant. Two Opus generations and the mini all cleared 22 on token budgets within 13 percent of each other, and the mini run twice against itself varied 24 percent, so the gap between models is smaller than the noise inside one of them. Pick the cheapest model that can drive the tools.

None of this is a new model of anything. It is AI Factory Economics run against a single Business Process Automation workload, with one unit of business output defined and everything denominated in it: cost per gate-verified repair. That framework warns that treating token consumption as a business metric is a trap, and I walked straight into it. The first version of this analysis was built on token efficiency, and measurement dissolved it. Tokens tracked nothing that mattered. The cost per finished repair tracked everything.

Scope it. Localized repair with an executable test, one task pool, one week of pricing. Three arms now have replicates, and they calibrate the rest: aggregate solve counts hold within a task of themselves, individual borderline tasks flip freely in both directions, and token budgets range from protocol-pinned in the loop, two runs within a fraction of a percent, to 24 percent apart in the harness. So trust the solve counts, treat any single task's verdict as weather, and read harness token comparisons only through noise that wide.

The obvious objection

Twenty-two tasks and mostly single runs. Isn't "the tier doesn't matter above the bar" exactly the kind of claim that noise produces?

The claim survives its own noise measurement, which is more than most tier comparisons attempt. Three arms were replicated: the free local loop held 17 twice, the paid loop moved 16 to 17, and the mini harness held 22 twice while its token count swung 24 percent. That swing is the calibration. The 13 percent token spread between a mini and two Opus generations sits inside it, so the honest statement is that the models are indistinguishable on consumption, and the fourfold cost gap is rate card. Solve counts replicate; individual tasks flip; token comparisons need error bars this wide. The entry says all three.

And the task pool is conditioned, deliberately. These are the 22 repairs a free local model could not clear on its first pass, which is the population that actually reaches an escalation decision. An unconditioned pool would flatter every paid tier by billing it for work the free rung would have absorbed.

Where it belongs

Where each layer belongs

LayerPlacement
Layer 0 · Compute
Compute & Network Fabric
The triage tier and the overflow buffer. Owned hardware runs the free loop at zero marginal cost, and the procurement finding is negative: the rentable GPU class cannot beat it, because agentic concurrency is VRAM-bound and the meter scales linearly with the work.
Retained
Layer 2B · Runtime
Application Runtime & Execution
The deterministic test gate and the loop orchestrator are code you run, not a model. The gate is also the yield instrument: it is what makes cost per verified repair computable, and minimum viable capability a measurement instead of a judgment call.
Retained
Layer 2C · Reasoning
Agentic Infrastructure — The Reasoning Plane
The determines-done authority never moves. Repair reasoning is delegated to whichever worker is cheapest above the instrument-literacy bar, and it is safe to shop that seat aggressively, in both directions, precisely because the authority stays here.
Retained
What it opened

The two questions this lab now knows to ask

How much work does a subscription seat's headroom actually hold?

The claim that idle Claude tokens do real work is structural here, priced off the rate card and the flat fee. The measured version runs an identical arm under subscription auth and meters cap consumption instead of dollars. That number would turn "spend headroom first" from a rule into a quantity, per seat, per week.

Where does the instrument-literacy bar actually sit?

Gemma 4 31B fails it and gpt-5.4-mini clears it, which brackets the bar but does not locate it. A descent through open-weight tiers inside the same harness would find the cheapest self-hostable model that still multiplies, and that model, not the frontier, is the interesting escalation target for anyone with idle VRAM.

Does the ladder survive losing localization?

Every task here arrived with its files named. Repository-scale fault localization is precisely the work an agentic harness should be best at and a constrained loop cannot attempt, so withholding localization should widen the harness's lead and may move the escalation boundary well below 17 of 22.

The bound

What it did not prove

  • Tier irrelevance above the bar is bounded by the noise measurement, not proved: 13 percent between models against 24 percent within one, on two runs. A wider replication could still separate them. What is excluded is any large token-efficiency gap of the kind the first draft of this analysis assumed.
  • The subscription rung's near-zero marginal cost is structural arithmetic, not a measurement. No arm ran under subscription auth, and the caps are not publicly token-denominated, so the crossover is a bound.
  • The task pool is the 22 repairs a specific free local model missed. That is the population an escalation decision actually sees, but it means every number here is conditioned on that model's failure profile, and the free rung's 17 of 22 does not transfer to other pools.
  • A one-task probe of Opus 5 extrapolated to a $19 to $20 arm; the measured arm cost $12.56. Single-task probes establish pricing and nothing else. This entry's per-arm figures are arm totals for exactly that reason.
In the author’s words

Notes from the author, Keith Townsend

Lab 011 ended by asking whether the agentic harness or the model tier carried its held-out fifth run, and this lab was built to answer that one question. It answered it and then kept moving. The harness carries it: a mini model in Claude Code matches two Opus generations exactly, and the tier decision collapsed into a rate-card decision. Both labs stand, and the pricing-regime footnote at the end of Lab 011 turned out to be the headline here.

Almost every number I believed at the start of this lab died during it. The token-efficiency thesis died when real usage capture showed the models consuming within noise of each other. A latency-tail claim died when the outlier that carried it failed to replicate. A cost claim flipped direction three times before both sides were measured on dedicated meters. The deterministic gate is the only reason any of those deaths were visible, which is the DCITL argument made by the lab's own errata.

Tied to the canon

Assessments at the time of the lab

NVIDIA AI PlatformLayer 0 · Compute
NVIDIA Strength — Silicon Authority · as assessed July 23, 2026 · current
NVIDIA AI PlatformLayer 2C · Reasoning
Runtime Governance Only — Not a Reasoning Plane · as assessed July 23, 2026 · current
AWS AI InfrastructureLayer 0 · Compute
Custom Silicon Full Stack · as assessed July 23, 2026 · current
Google Cloud AI InfrastructureLayer 0 · Compute
TPU + GPU Full Stack · as assessed July 23, 2026 · current
Google Cloud AI InfrastructureLayer 2C · Reasoning
Productized Placement · as assessed July 23, 2026 · current
How it was built

Method and disclosure

Self-funded editorial. The harness is loopcontrolbench, published open under the MIT license at github.com/kltownsend/loopcontrolbench: the reproduce-or-drop gate, the sandboxed runner, and the arm definitions, including the Claude Code driver and the translation-proxy configuration added for this lab. The vendor assessments and the layer placements this evidence feeds stay proprietary.

This lab tests an existing model rather than proposing one. The AI Factory Economics Framework was published at thectoadvisor.com/blog/2026/01/07/from-ai-factory-metaphor-to-cio-grade-economics/, and every cost figure here is denominated in that framework's unit of business output: one gate-verified repair. The framework names token fixation as a failure mode; the first version of this analysis committed it, and the correction is reported in the writeup rather than quietly removed.