# Lab 012 — Raw lab detail: Buy the harness, not the tier

Working notes behind the published entry. This lab's findings died and were replaced several
times on the way to publication; the deaths are recorded here because the gate is what caught
every one of them.

## The one question

Lab 011 ended with a held-out run it could not attribute: a top model inside Claude Code
cleared all 22 misses, but that swapped the harness and the model at once. This lab holds
Claude Code fixed and moves the model, from a free local Gemma to two generations of Opus,
and prices every rung an architect could actually buy.

## The instrument

The 22 tasks Lab 011's free local loop first missed, each certified by a reproduce-or-drop
gate. Fourteen arms across two apparatuses: the constrained repair loop, and Claude Code
driven headless through a LiteLLM translation proxy. Three arms replicated. Cost capture is
the method contribution: a dedicated API key isolating one arm's dashboard line, and per-task
usage envelopes on the native Anthropic path. No apportioned dollar figure survived to
publication.

## The measured spine

- Free local 31B loop: 17 and 17 of 22 across two runs. Paid mini in the same loop: 16 and
  17. Dead parity; the tier bought no separation in the same apparatus.
- gpt-5.4-mini in Claude Code: 22 of 22, twice, $3.14 measured. Opus 4.8: 22 of 22, $13.79.
  Opus 5: 22 of 22, $12.56.
- Token budgets across that fourfold price range: 10.38M, 10.96M, 9.66M. A 13 percent spread
  against the mini's own 24 percent run-to-run variance. The harness sets the token budget;
  the tier multiplies the rate card.
- Escalating only the five the gate named: $1.12. The residual five are 23 percent of tasks
  and roughly half the bill; hard tasks are expensive in proportion to being hard.
- Instrument literacy gates substitution: Gemma 31B went 17 in the loop to 11 in the harness.
  The harness multiplies a model that can operate its instruments and taxes one that cannot.

## Findings that died in measurement (kept deliberately)

1. Token efficiency. The draft's original headline had Opus clearing the pool in a fraction
   of the mini's tokens, built on a hand-recorded 1.28M from Lab 011. Real usage capture:
   10.96M, off by 8.6x. The thesis dissolved; the models consume within noise of each other.
2. The latency tail. One 26-minute outlier carried a "the premium buys predictability" claim.
   The replicate maxed at 4 minutes. Run variance, not model property.
3. The cost direction itself flipped three times before both sides were measured on dedicated
   meters. Token fixation is named as a trap by the AI Factory Economics framework this lab
   tests; the first draft committed it, and the correction ships in the writeup.

## The procurement rulings

Renting the available GPU class to self-host the escalation rung loses on cost and wall clock
at once: agentic sessions park context in VRAM while barely touching compute, so concurrency
is bought with memory and the meter scales linearly with the work. Per stream per hour, the
owned box holds most of a rented fleet for nothing. Scale out with the smallest instance,
never up; for this rung, do not scale at all. Rent the tokens.

## The ladder, and its two directions

Spend committed capacity first at every rung it exists: owned idle hardware, unused
cloud-spend commitment, unused subscription headroom. One category, three altitudes.
Escalate what the gate hands back to the cheapest instrument-literate model. Route overflow
back down when the budget binds instead of the clock; the gate makes both directions safe
because a verified repair is verified regardless of which rung produced it.

## Operational record (mistakes included)

- A first Opus attempt died at task 14 on credit exhaustion: $5.45 of errored sessions,
  deleted and re-run. Check the resource a metered run spends against before launching.
- A g6e.xlarge idled 12 hours 42 minutes (~$24) because a stale note stood in for a check
  expired credentials could not perform. "Unverified" is the only honest answer when the
  console is unreachable.
- A one-task probe of Opus 5 extrapolated to a $19-20 arm; the arm cost $12.56. Probes
  establish pricing and nothing else.
- The mistakes stay in the published bill.

## Spend

About $121 total: $73.75 of rented infrastructure and $47 and change in metered tokens,
including the waste itemized above. Both 22-of-22 configurations landed in single digits of
dollars; the instrument that told them apart cost roughly one more arm on top.

## What this hands labs 013 and 014

Two questions this lab could not answer with a saturated pool. The self-hosted floor: what is
the entry-level open-weight model that clears the instrument-literacy bar Gemma failed. And
the tier boundary: mini and Opus tied at 22 of 22 here, so the difficulty line where the 4x
rate card starts buying solves has to be found on a harder pool, mined for it. The task
difficulty rating built from this lab's fourteen arms is the selection instrument for both.
