# Buy the harness, not the tier

> Lab 012 · Editorial lab · Status: published  
> Published by The CTO Advisor LLC · Layer2C Labs

**Question:** An architect escalating repair work has rungs available: a free local loop, a paid mid-tier model in that same loop, then a frontier model inside an agentic harness. I priced all of them. Same 22 certified bug-fix repairs, same deterministic test gate, six models from a local Gemma to two generations of Opus, run both as a constrained loop and inside Claude Code. The free local loop clears 17. The paid mid-tier loop clears 16 and 17 across two runs, no better. The harness is what closes the rest, and it does that for a mini model as readily as a frontier one. So which rung is actually load-bearing, and what does each one cost?

**Load:** The 22 tasks Lab 011's free local loop first missed, each a real bug-fix pull request shipping its own test suite as an unfalsifiable pass/fail gate. Fourteen arms across two apparatuses: the constrained repair loop from Lab 011, and Claude Code driven headless through a translation proxy. Workers ranged from Gemma 4 26B and 31B on owned hardware, through hosted gpt-5-mini and gpt-5.4-mini, to Claude Opus 4.8 and Opus 5. Three arms were replicated. Every cost figure is directly measured: the mini arm from a dedicated API key's dashboard line, the Opus arms from per-task usage envelopes on the native API path. Total spend: about $121. $73.75 of rented infrastructure and $47 and change in metered tokens, including $5.45 killed by a mid-run credit exhaustion and roughly $24 of GPU left idling overnight. The mistakes stay in the bill because a reproducer pays for theirs too. The Spark's triage rung added nothing to any meter.

## Executive Summary

The pitch under test: routing repair work to a cheaper model saves money, and the counter-pitch that cheap models flail and cost more by the end. Both failed. Across 14 arms on the same 22 gate-verified repairs, capability entered through the apparatus, not the price column, and the cost of finished work tracked the rate card alone.

The ladder, priced per verified repair: free local triage clears 17 of 22 at roughly zero. A paid mid-tier model in the same loop clears 16 and 17 across two runs, no separation from free. The same mini inside Claude Code clears all 22 for $3.14. Two Opus generations clear the identical 22 for $13.79 and $12.56. Escalating only the five the gate names costs $1.12.

Renting GPU to self-host the escalation rung fails on both axes at once. Agentic sessions park context in VRAM while barely touching compute, so concurrency is bought with memory, and the box already on the bench holds most of a rented fleet's sessions for nothing. The rented arm cost more than the API and solved fewer than half the tasks.

The transferable rule is a dispatcher, not a ladder. Spend committed capacity first, at every rung it exists: idle hardware, unused cloud commitment, unused subscription headroom. Escalate to metered capability only for what the gate hands back, buy the minimum tier that clears the bar, and route overflow back down when the budget binds instead of the clock.

## DAPM Table — Authority Verdict

| Layer | Placement |
| --- | --- |
| layer0 | Retained |
| layer2b | Retained |
| layer2c | Retained |

## Seam Map — Readiness

| Function | Readiness |
| --- | --- |

## Detailed Writeup

Start on capacity you already carry. A Gemma 4 31B in a constrained loop clears 17 of the 22 on deterministic test feedback alone, and the gate hands back the exact five it could not solve. The earlier Xeon lab priced that rung: about nine cents per verified repair on rented Granite Rapids, six at the measured Savings Plan rate, and approaching zero on owned idle hardware or inside a commitment you already bought and did not use. This is not a story about owning a Spark. Unused cloud-spend commitment is idle capacity too.

The next rung up is the one that doesn't work. A paid mid-tier model in that same constrained loop scored 16 of 22, and 17 on a second run, against the free local model's 17 and 17. Two runs each, dead parity. Buying a better model without changing the apparatus around it bought no separation from free. That is the result that reframes the ladder, because it means the money is not in the tier.

It's in the harness. The same mini model that plateaued at 16 and 17 in the loop clears all 22 inside Claude Code. Six tasks recovered from apparatus, not parameters. That is where capability actually enters, and it is not exotic: read the repository, run the tests, read the failure, try again, with tools instead of a fixed protocol. Escalating just the residual five through it costs $1.12 against $12.56 for pushing all 22 through a frontier model, an eleven-fold spread for the same finished work. And even that $1.12 is a spot price. A flat-rate subscription seat with weekly headroom left over is idle capacity at this rung, the same way the Xeon lab's spare cycles were idle capacity at the one below. Idle Claude tokens can do real work too, and the marginal cost of the residual five inside an already-paid seat rounds to zero until the cap binds.

So should you buy the GPU class you can actually get? I priced it, and no. Agentic sessions are bought with memory, not compute: each one parks its context for its whole life while barely touching the die, so the honest unit is dollars per concurrent session per hour. Four entry instances give about nineteen sessions for $7.44 an hour. The box already on my bench holds fifteen for nothing. Renting the liquid class buys four more sessions at seven dollars an hour, which is not a purchase, it is a rounding error with a meter attached. The tier above is worse on the same axis and quota-locked besides, and it sells memory bandwidth against a per-task latency that is serial and irreducible. If you do rent, scale out with the smallest instance rather than up: consolidated multi-GPU boxes charge you for interconnect that independent sessions never touch.

What separates the rungs is what a retry costs, and that decides which model belongs on each. On idle capacity a retry is free, so a variance-prone cheap model is not a liability, it is a method: let it grind and let the gate certify each attempt. The Xeon lab audited 585 scored attempts and found no validator pass that failed its held-out checks, which is what makes grinding safe rather than reckless. On a metered API the arithmetic inverts. Every retry bills, so what you are actually buying up the ladder is first-pass success. That is the same property as instrument literacy, seen from the invoice.

One constraint keeps you honest, and it is why the escalation rung has to be bought rather than substituted. You cannot simply wrap the free local model in the harness to skip the paid step: Gemma 4 31B goes from 17 down to 11 inside Claude Code. The harness multiplies a model that can operate its instruments and taxes one that cannot. Instrument literacy is its own capability and it does not track parameter count. Above that bar the tier is close to irrelevant. Two Opus generations and the mini all cleared 22 on token budgets within 13 percent of each other, and the mini run twice against itself varied 24 percent, so the gap between models is smaller than the noise inside one of them. Pick the cheapest model that can drive the tools.

None of this is a new model of anything. It is AI Factory Economics run against a single Business Process Automation workload, with one unit of business output defined and everything denominated in it: cost per gate-verified repair. That framework warns that treating token consumption as a business metric is a trap, and I walked straight into it. The first version of this analysis was built on token efficiency, and measurement dissolved it. Tokens tracked nothing that mattered. The cost per finished repair tracked everything.

Scope it. Localized repair with an executable test, one task pool, one week of pricing. Three arms now have replicates, and they calibrate the rest: aggregate solve counts hold within a task of themselves, individual borderline tasks flip freely in both directions, and token budgets range from protocol-pinned in the loop, two runs within a fraction of a percent, to 24 percent apart in the harness. So trust the solve counts, treat any single task's verdict as weather, and read harness token comparisons only through noise that wide.

## How It Abstracts

The shape generalizes to any Business Process Automation workload with a deterministic acceptance test. Triage on capacity you already carry, escalate only what the gate hands back, and spend committed capacity at every rung before buying metered capability: idle hardware, unused cloud-spend commitment, and unused subscription headroom are one category at three altitudes. Only when all of it is spent do you buy, and then the cheapest input that clears the yield bar.

The ladder also runs both directions. Escalate when the clock is the binding constraint. Route work back down when the budget is: an exhausted allowance or a capped subscription week sends overflow to hardware that costs nothing to keep busy, and the free rung becomes the buffer instead of the entry point. The gate is what makes both directions safe. A verified repair is a verified repair regardless of which rung produced it, so down-routing degrades only time and yield, both of which you can see, never quality silently.

The transferable instrument is the denominator. Cost per verified unit of output is what made every comparison here legible, and it is the one thing an architect has to define before any of this arithmetic runs. Without it, tier debates are unfalsifiable.

## Assessments at the Time of the Lab

| Vendor | Layer | Grade | As assessed |
| --- | --- | --- | --- |
| NVIDIA AI Platform | Layer 0 · Compute | NVIDIA Strength — Silicon Authority | July 23, 2026 |
| NVIDIA AI Platform | Layer 2C · Reasoning | Runtime Governance Only — Not a Reasoning Plane | July 23, 2026 |
| AWS AI Infrastructure | Layer 0 · Compute | Custom Silicon Full Stack | July 23, 2026 |
| Google Cloud AI Infrastructure | Layer 0 · Compute | TPU + GPU Full Stack | July 23, 2026 |
| Google Cloud AI Infrastructure | Layer 2C · Reasoning | Productized Placement | July 23, 2026 |

## Method and Disclosure

Self-funded editorial. The harness is loopcontrolbench, published open under the MIT license at github.com/kltownsend/loopcontrolbench: the reproduce-or-drop gate, the sandboxed runner, and the arm definitions, including the Claude Code driver and the translation-proxy configuration added for this lab. The vendor assessments and the layer placements this evidence feeds stay proprietary.

This lab tests an existing model rather than proposing one. The AI Factory Economics Framework was published at thectoadvisor.com/blog/2026/01/07/from-ai-factory-metaphor-to-cio-grade-economics/, and every cost figure here is denominated in that framework's unit of business output: one gate-verified repair. The framework names token fixation as a failure mode; the first version of this analysis committed it, and the correction is reported in the writeup rather than quietly removed.

---
*Layer2C Labs · The CTO Advisor LLC · labs.layer2c.com*
