{
  "slug": "harness-or-tier",
  "lab_number": 12,
  "title": "Buy the harness, not the tier",
  "type": "editorial",
  "status": "published",
  "sponsor": null,
  "author": "Keith Townsend",
  "date": "July 26, 2026",
  "date_iso": "2026-07-26",
  "layers": [
    "layer0",
    "layer2b",
    "layer2c"
  ],
  "axes": [
    {
      "instrument": "4plus1",
      "layers": [
        "layer0",
        "layer2b",
        "layer2c"
      ],
      "has_dapm_table": true
    }
  ],
  "supersedes": [],
  "superseded_by": null,
  "corrects": [],
  "corrected_by": [],
  "vendors": [
    {
      "key": "nvidia",
      "role": "hw"
    },
    {
      "key": "google",
      "role": "model"
    },
    {
      "key": "openai",
      "role": "model"
    },
    {
      "key": "anthropic",
      "role": "model"
    },
    {
      "key": "aws",
      "role": "cloud"
    },
    {
      "key": "gcp",
      "role": "cloud"
    }
  ],
  "themes": [
    "loop-control",
    "ai-factory-economics",
    "validator-authority",
    "agentic-repair",
    "recovered-capacity"
  ],
  "finding": "Open weights on capacity I already carry clear 17 of 22 repairs for nothing, and the gate names the five they miss. Renting more hardware for the rest is slower and dearer than the API. The minimum that finishes the job is a mini model in an agentic harness, at $3.14. Two Opus generations finish the same 22 for four times that.",
  "question": "An architect escalating repair work has rungs available: a free local loop, a paid mid-tier model in that same loop, then a frontier model inside an agentic harness. I priced all of them. Same 22 certified bug-fix repairs, same deterministic test gate, six models from a local Gemma to two generations of Opus, run both as a constrained loop and inside Claude Code. The free local loop clears 17. The paid mid-tier loop clears 16 and 17 across two runs, no better. The harness is what closes the rest, and it does that for a mini model as readily as a frontier one. So which rung is actually load-bearing, and what does each one cost?",
  "load": "The 22 tasks Lab 011's free local loop first missed, each a real bug-fix pull request shipping its own test suite as an unfalsifiable pass/fail gate. Fourteen arms across two apparatuses: the constrained repair loop from Lab 011, and Claude Code driven headless through a translation proxy. Workers ranged from Gemma 4 26B and 31B on owned hardware, through hosted gpt-5-mini and gpt-5.4-mini, to Claude Opus 4.8 and Opus 5. Three arms were replicated. Every cost figure is directly measured: the mini arm from a dedicated API key's dashboard line, the Opus arms from per-task usage envelopes on the native API path. Total spend: about $121. $73.75 of rented infrastructure and $47 and change in metered tokens, including $5.45 killed by a mid-run credit exhaustion and roughly $24 of GPU left idling overnight. The mistakes stay in the bill because a reproducer pays for theirs too. The Spark's triage rung added nothing to any meter.",
  "verdict": {
    "scope": "Measured on localized repair with an executable test, one task pool of 22, one week of vendor pricing. The rulings are about where capability enters an agentic stack and what each rung of an escalation ladder costs per verified unit of output. The open edges are repository-scale debugging without localization, domains without an executable evaluator, and variance bounds beyond two runs per arm.",
    "independence": "Self-funded editorial. No vendor paid for this answer, reviewed it, or saw it before publication. The free local worker is Gemma, a Google model, run on owned NVIDIA hardware; the paid arms were metered OpenAI and Anthropic calls, and the rented-GPU arms ran on AWS. Disclosure: Google Cloud is a client of The CTO Advisor LLC; NVIDIA, OpenAI, Anthropic, and AWS are not. The ruling is the author's alone.",
    "calls": [
      {
        "kind": "do",
        "text": "triage on capacity you already carry. The free local loop clears three quarters of this work at zero marginal cost, and idle capacity comes in three forms: owned hardware, unused cloud-spend commitment, and unused subscription headroom."
      },
      {
        "kind": "do",
        "text": "escalate what the gate hands back into an agentic harness, and put the cheapest instrument-literate model in it. The mini cleared all 22 for $3.14; two Opus generations cleared the same 22 for four times that."
      },
      {
        "kind": "do",
        "text": "run the ladder in both directions. Escalate when the clock binds. Route overflow back to idle hardware when the budget binds. The gate makes both directions safe because a verified repair is verified regardless of which rung produced it."
      },
      {
        "kind": "dont",
        "text": "buy a better model for the constrained loop. The paid mid-tier loop scored 16 and 17 across two runs against the free local model's 17 and 17. Same apparatus, no separation from free."
      },
      {
        "kind": "dont",
        "text": "rent GPU to self-host the escalation rung. Agentic sessions are VRAM-bound, not compute-bound, so the rented class costs more per verified repair than the API and delivers less than half the capability. You cannot rent your way to efficiency on this class of model."
      },
      {
        "kind": "dont",
        "text": "read token counts as a business metric. Three models spanning a fourfold price range landed within 13 percent of each other on tokens, and one model run twice against itself varied 24 percent. Cost per verified repair is the number that decides anything."
      }
    ]
  },
  "objection": {
    "q": "Twenty-two tasks and mostly single runs. Isn't \"the tier doesn't matter above the bar\" exactly the kind of claim that noise produces?",
    "a": [
      "The claim survives its own noise measurement, which is more than most tier comparisons attempt. Three arms were replicated: the free local loop held 17 twice, the paid loop moved 16 to 17, and the mini harness held 22 twice while its token count swung 24 percent. That swing is the calibration. The 13 percent token spread between a mini and two Opus generations sits inside it, so the honest statement is that the models are indistinguishable on consumption, and the fourfold cost gap is rate card. Solve counts replicate; individual tasks flip; token comparisons need error bars this wide. The entry says all three.",
      "And the task pool is conditioned, deliberately. These are the 22 repairs a free local model could not clear on its first pass, which is the population that actually reaches an escalation decision. An unconditioned pool would flatter every paid tier by billing it for work the free rung would have absorbed."
    ]
  },
  "numbers": [
    {
      "label": "Configurations that cleared all 22",
      "value": "4",
      "note": "gpt-5.4-mini twice, Opus 4.8, Opus 5, all inside the same agentic harness"
    },
    {
      "label": "Cost per verified repair, mini harness vs Opus",
      "value": "$0.14 vs $0.63",
      "note": "$3.14 against $13.79 for identical finished work; both directly measured"
    },
    {
      "label": "Free local loop, two runs",
      "value": "17 and 17 of 22",
      "note": "paid mid-tier in the same loop: 16 and 17, no separation from free"
    },
    {
      "label": "Token spread across a 4x price range",
      "value": "13%",
      "note": "inside the mini's own 24% run-to-run variance; the harness sets the token budget, the tier multiplies the rate"
    },
    {
      "label": "Rented-GPU self-host of the escalation rung",
      "value": "9 of 22",
      "note": "at $4-6 of machine time: more than the API, for under half the capability"
    },
    {
      "label": "Total spend, every meter, mistakes included",
      "value": "~$121",
      "note": "$73.75 rented infrastructure, $47 and change in metered tokens, including $5.45 lost to a mid-run credit exhaustion and ~$24 of idle GPU; a reproducer pays for their failures too"
    }
  ],
  "opened": [
    {
      "q": "How much work does a subscription seat's headroom actually hold?",
      "a": "The claim that idle Claude tokens do real work is structural here, priced off the rate card and the flat fee. The measured version runs an identical arm under subscription auth and meters cap consumption instead of dollars. That number would turn \"spend headroom first\" from a rule into a quantity, per seat, per week."
    },
    {
      "q": "Where does the instrument-literacy bar actually sit?",
      "a": "Gemma 4 31B fails it and gpt-5.4-mini clears it, which brackets the bar but does not locate it. A descent through open-weight tiers inside the same harness would find the cheapest self-hostable model that still multiplies, and that model, not the frontier, is the interesting escalation target for anyone with idle VRAM."
    },
    {
      "q": "Does the ladder survive losing localization?",
      "a": "Every task here arrived with its files named. Repository-scale fault localization is precisely the work an agentic harness should be best at and a constrained loop cannot attempt, so withholding localization should widen the harness's lead and may move the escalation boundary well below 17 of 22."
    }
  ],
  "not_proved": [
    "Tier irrelevance above the bar is bounded by the noise measurement, not proved: 13 percent between models against 24 percent within one, on two runs. A wider replication could still separate them. What is excluded is any large token-efficiency gap of the kind the first draft of this analysis assumed.",
    "The subscription rung's near-zero marginal cost is structural arithmetic, not a measurement. No arm ran under subscription auth, and the caps are not publicly token-denominated, so the crossover is a bound.",
    "The task pool is the 22 repairs a specific free local model missed. That is the population an escalation decision actually sees, but it means every number here is conditioned on that model's failure profile, and the free rung's 17 of 22 does not transfer to other pools.",
    "A one-task probe of Opus 5 extrapolated to a $19 to $20 arm; the measured arm cost $12.56. Single-task probes establish pricing and nothing else. This entry's per-arm figures are arm totals for exactly that reason."
  ],
  "executive_summary": [
    "The pitch under test: routing repair work to a cheaper model saves money, and the counter-pitch that cheap models flail and cost more by the end. Both failed. Across 14 arms on the same 22 gate-verified repairs, capability entered through the apparatus, not the price column, and the cost of finished work tracked the rate card alone.",
    "The ladder, priced per verified repair: free local triage clears 17 of 22 at roughly zero. A paid mid-tier model in the same loop clears 16 and 17 across two runs, no separation from free. The same mini inside Claude Code clears all 22 for $3.14. Two Opus generations clear the identical 22 for $13.79 and $12.56. Escalating only the five the gate names costs $1.12.",
    "Renting GPU to self-host the escalation rung fails on both axes at once. Agentic sessions park context in VRAM while barely touching compute, so concurrency is bought with memory, and the box already on the bench holds most of a rented fleet's sessions for nothing. The rented arm cost more than the API and solved fewer than half the tasks.",
    "The transferable rule is a dispatcher, not a ladder. Spend committed capacity first, at every rung it exists: idle hardware, unused cloud commitment, unused subscription headroom. Escalate to metered capability only for what the gate hands back, buy the minimum tier that clears the bar, and route overflow back down when the budget binds instead of the clock."
  ],
  "dapm": [
    {
      "instrument": "4plus1",
      "layer": "layer0",
      "placement": "Retained",
      "note": "The triage tier and the overflow buffer. Owned hardware runs the free loop at zero marginal cost, and the procurement finding is negative: the rentable GPU class cannot beat it, because agentic concurrency is VRAM-bound and the meter scales linearly with the work."
    },
    {
      "instrument": "4plus1",
      "layer": "layer2b",
      "placement": "Retained",
      "note": "The deterministic test gate and the loop orchestrator are code you run, not a model. The gate is also the yield instrument: it is what makes cost per verified repair computable, and minimum viable capability a measurement instead of a judgment call."
    },
    {
      "instrument": "4plus1",
      "layer": "layer2c",
      "placement": "Retained",
      "note": "The determines-done authority never moves. Repair reasoning is delegated to whichever worker is cheapest above the instrument-literacy bar, and it is safe to shop that seat aggressively, in both directions, precisely because the authority stays here."
    }
  ],
  "seam_map": [],
  "writeup": [
    "Start on capacity you already carry. A Gemma 4 31B in a constrained loop clears 17 of the 22 on deterministic test feedback alone, and the gate hands back the exact five it could not solve. The earlier Xeon lab priced that rung: about nine cents per verified repair on rented Granite Rapids, six at the measured Savings Plan rate, and approaching zero on owned idle hardware or inside a commitment you already bought and did not use. This is not a story about owning a Spark. Unused cloud-spend commitment is idle capacity too.",
    "The next rung up is the one that doesn't work. A paid mid-tier model in that same constrained loop scored 16 of 22, and 17 on a second run, against the free local model's 17 and 17. Two runs each, dead parity. Buying a better model without changing the apparatus around it bought no separation from free. That is the result that reframes the ladder, because it means the money is not in the tier.",
    "It's in the harness. The same mini model that plateaued at 16 and 17 in the loop clears all 22 inside Claude Code. Six tasks recovered from apparatus, not parameters. That is where capability actually enters, and it is not exotic: read the repository, run the tests, read the failure, try again, with tools instead of a fixed protocol. Escalating just the residual five through it costs $1.12 against $12.56 for pushing all 22 through a frontier model, an eleven-fold spread for the same finished work. And even that $1.12 is a spot price. A flat-rate subscription seat with weekly headroom left over is idle capacity at this rung, the same way the Xeon lab's spare cycles were idle capacity at the one below. Idle Claude tokens can do real work too, and the marginal cost of the residual five inside an already-paid seat rounds to zero until the cap binds.",
    "So should you buy the GPU class you can actually get? I priced it, and no. Agentic sessions are bought with memory, not compute: each one parks its context for its whole life while barely touching the die, so the honest unit is dollars per concurrent session per hour. Four entry instances give about nineteen sessions for $7.44 an hour. The box already on my bench holds fifteen for nothing. Renting the liquid class buys four more sessions at seven dollars an hour, which is not a purchase, it is a rounding error with a meter attached. The tier above is worse on the same axis and quota-locked besides, and it sells memory bandwidth against a per-task latency that is serial and irreducible. If you do rent, scale out with the smallest instance rather than up: consolidated multi-GPU boxes charge you for interconnect that independent sessions never touch.",
    "What separates the rungs is what a retry costs, and that decides which model belongs on each. On idle capacity a retry is free, so a variance-prone cheap model is not a liability, it is a method: let it grind and let the gate certify each attempt. The Xeon lab audited 585 scored attempts and found no validator pass that failed its held-out checks, which is what makes grinding safe rather than reckless. On a metered API the arithmetic inverts. Every retry bills, so what you are actually buying up the ladder is first-pass success. That is the same property as instrument literacy, seen from the invoice.",
    "One constraint keeps you honest, and it is why the escalation rung has to be bought rather than substituted. You cannot simply wrap the free local model in the harness to skip the paid step: Gemma 4 31B goes from 17 down to 11 inside Claude Code. The harness multiplies a model that can operate its instruments and taxes one that cannot. Instrument literacy is its own capability and it does not track parameter count. Above that bar the tier is close to irrelevant. Two Opus generations and the mini all cleared 22 on token budgets within 13 percent of each other, and the mini run twice against itself varied 24 percent, so the gap between models is smaller than the noise inside one of them. Pick the cheapest model that can drive the tools.",
    "None of this is a new model of anything. It is AI Factory Economics run against a single Business Process Automation workload, with one unit of business output defined and everything denominated in it: cost per gate-verified repair. That framework warns that treating token consumption as a business metric is a trap, and I walked straight into it. The first version of this analysis was built on token efficiency, and measurement dissolved it. Tokens tracked nothing that mattered. The cost per finished repair tracked everything.",
    "Scope it. Localized repair with an executable test, one task pool, one week of pricing. Three arms now have replicates, and they calibrate the rest: aggregate solve counts hold within a task of themselves, individual borderline tasks flip freely in both directions, and token budgets range from protocol-pinned in the loop, two runs within a fraction of a percent, to 24 percent apart in the harness. So trust the solve counts, treat any single task's verdict as weather, and read harness token comparisons only through noise that wide."
  ],
  "abstraction": [
    "The shape generalizes to any Business Process Automation workload with a deterministic acceptance test. Triage on capacity you already carry, escalate only what the gate hands back, and spend committed capacity at every rung before buying metered capability: idle hardware, unused cloud-spend commitment, and unused subscription headroom are one category at three altitudes. Only when all of it is spent do you buy, and then the cheapest input that clears the yield bar.",
    "The ladder also runs both directions. Escalate when the clock is the binding constraint. Route work back down when the budget is: an exhausted allowance or a capped subscription week sends overflow to hardware that costs nothing to keep busy, and the free rung becomes the buffer instead of the entry point. The gate is what makes both directions safe. A verified repair is a verified repair regardless of which rung produced it, so down-routing degrades only time and yield, both of which you can see, never quality silently.",
    "The transferable instrument is the denominator. Cost per verified unit of output is what made every comparison here legible, and it is the one thing an architect has to define before any of this arithmetic runs. Without it, tier debates are unfalsifiable."
  ],
  "method": [
    "Self-funded editorial. The harness is loopcontrolbench, published open under the MIT license at github.com/kltownsend/loopcontrolbench: the reproduce-or-drop gate, the sandboxed runner, and the arm definitions, including the Claude Code driver and the translation-proxy configuration added for this lab. The vendor assessments and the layer placements this evidence feeds stay proprietary.",
    "This lab tests an existing model rather than proposing one. The AI Factory Economics Framework was published at thectoadvisor.com/blog/2026/01/07/from-ai-factory-metaphor-to-cio-grade-economics/, and every cost figure here is denominated in that framework's unit of business output: one gate-verified repair. The framework names token fixation as a failure mode; the first version of this analysis committed it, and the correction is reported in the writeup rather than quietly removed."
  ],
  "assessments": [
    {
      "instrument": "4plus1",
      "vendor": "NVIDIA AI Platform",
      "layer": "Layer 0 · Compute",
      "grade": "NVIDIA Strength — Silicon Authority",
      "assessed_date": "July 23, 2026",
      "version": "v1.5 - Idle-Silicon Economics Reconciliation",
      "source_url": "https://layer2c.com/assessment/nvidia",
      "snapshot_date": "2026-07-26"
    },
    {
      "instrument": "4plus1",
      "vendor": "NVIDIA AI Platform",
      "layer": "Layer 2C · Reasoning",
      "grade": "Runtime Governance Only — Not a Reasoning Plane",
      "assessed_date": "July 23, 2026",
      "version": "v1.5 - Idle-Silicon Economics Reconciliation",
      "source_url": "https://layer2c.com/assessment/nvidia",
      "snapshot_date": "2026-07-26"
    },
    {
      "instrument": "4plus1",
      "vendor": "AWS AI Infrastructure",
      "layer": "Layer 0 · Compute",
      "grade": "Custom Silicon Full Stack",
      "assessed_date": "July 23, 2026",
      "version": "v1.5 - Inference-Interface Notes",
      "source_url": "https://layer2c.com/assessment/aws",
      "snapshot_date": "2026-07-26"
    },
    {
      "instrument": "4plus1",
      "vendor": "Google Cloud AI Infrastructure",
      "layer": "Layer 0 · Compute",
      "grade": "TPU + GPU Full Stack",
      "assessed_date": "July 23, 2026",
      "version": "v1.6 - TPU BYO-Serving Reconciliation",
      "source_url": "https://layer2c.com/assessment/gcp",
      "snapshot_date": "2026-07-26"
    },
    {
      "instrument": "4plus1",
      "vendor": "Google Cloud AI Infrastructure",
      "layer": "Layer 2C · Reasoning",
      "grade": "Productized Placement",
      "assessed_date": "July 23, 2026",
      "version": "v1.6 - TPU BYO-Serving Reconciliation",
      "source_url": "https://layer2c.com/assessment/gcp",
      "snapshot_date": "2026-07-26"
    }
  ],
  "publisher": "The CTO Advisor LLC",
  "program": "Layer2C Labs",
  "url": "https://labs.layer2c.com/labs/harness-or-tier",
  "markdown_url": "https://labs.layer2c.com/labs/harness-or-tier.md",
  "json_url": "https://labs.layer2c.com/labs/harness-or-tier.json"
}
