Recovered capacity is real, and it fails honest
By Keith Townsend · 2026-07-16
Enterprises commonly carry idle owned hardware or unused cloud-spend commitments above their operating baseline: headroom producing nothing between peaks. This lab asked one question of that idle capacity: can a smaller model clear real, verifiable work on it, judged by the same deterministic validator that judged the big model? Across three models, four quantization tiers, and two substrates, the answer came back yes, with the load-bearing detail attached: in 585 scored attempts, the audit found no validator pass that failed the available held-out checks. The work that cleared was verified as far as the instrument can see. The work that failed went to a human queue. No token meter ran.
Scoped to falsifiable batch work: tasks with a deterministic validator, here real-library bug fixes judged by executable tests. The safety claim extends exactly as far as the validator’s detection power; deterministic does not mean complete, and a deterministic gate can reproducibly admit a defect its tests do not detect. The ruling answers the one question; the boundary numbers are this environment’s coordinates, not universal constants. The size and shape of the models and compute another team needs is their sizing exercise, and this page ships the instrument and method to run it. Two adjacent cells were deliberately not run and are named in the bound: the 26B MoE at Q4 and the E2B size rung. The ruling does not need them, and asserting cells you chose not to measure is how labs drift into marketing.
Do Put committed idle capacity to work on falsifiable batches behind a deterministic validator. Measured yield: 51 to 54% of attempts cleared as verified passes among the configurations above the measured boundary (numerators on the card), with the Lab 3 baseline at 47.8% on its own 29-task coverage. On the rented Granite Rapids shape, the full 117-attempt protocol produced 63 verified bug fixes in 2.24 hours: about nine cents of on-demand compute per verified fix, six cents at the measured one-year Savings Plan rate, and an incremental bill that can approach zero inside an otherwise-unused spend commitment.
Do Trust the gate, not the model. Zero confirmed false passes in 585 scored attempts across every configuration measured. Failures routed to the human queue or died at the output gate; no confirmed false success cleared the measured gate, and that claim extends exactly as far as the validator and its held-out evidence detect. The failure mode of this architecture is a person looks at it, not a token meter running.
Do Size down with a controlled descent. In this environment the measured model boundary landed at a 7.7GB artifact: Gemma 4 12B at Q4_K_M held the identical pass set as its bf16 original across three quantization tiers, measured on the Spark control. Substrate survival was proven at bf16, so the boundary config on rented Xeon rests on two measured edges rather than a measured cell; that bound is stated, not hidden. One rung lower broke, cleanly and honestly. Your boundary will differ; the screen-then-confirm ladder that finds it takes an afternoon.
Don’t Do not require the model to be deterministic, and do not read variance as defect. The deterministic component of this architecture is the validator. A variance-prone worker behind a deterministic gate fails closed against the defects the gate is built to detect, and on idle cycles its retries become mining: every additional pass it flickers into is checked against the held-out evidence before it counts.
Disclosure: Intel is a client of this practice, and the lab corresponded with Intel about serving configurations during the program; the exchange is paraphrased where it appears and no vendor previewed or funded any of it. The bench ran on $28.32 of self-funded cloud spend plus owned hardware. No vendor paid for this answer.
Most enterprises carry headroom that produces nothing: owned hardware idling between peaks, or a cloud-spend commitment running under coverage. This lab asked one question of that idle capacity. Can a smaller model clear real, verifiable work on it, judged by the same deterministic validator that judged the big model? Across three models, four quantization tiers, and two substrates, the answer came back yes, and it came back with the load-bearing detail attached: in 585 scored attempts, the audit found no validator pass that failed the available held-out checks. The work that cleared was verified as far as the instrument can see. The work that failed went to a human queue. No token meter ran.
The instrument is the credibility, so it goes first. This is the Lab 3 harness, unmodified and pin-verified: 39 coding tasks, most reconstructed from real bug fixes in real Python libraries, httpx, requests, dateutil, more-itertools, arrow, and click among them, each pinned to a repo and commit with the historical bug present. Detection is layered on purpose. Feature tests verify the fix works, regression tests catch collateral damage, hidden reference tests the model never sees catch gaming, and a static gate kills malformed output before any code runs. The audit that verified the validator also X-rayed it: one weak test, two tasks without hidden references, six structurally unpassable tasks, all found by running the same audit against the original baseline and carried like-for-like on both sides of every comparison. A lab that won't X-ray its own instrument has no business grading anyone else's. The workers were Gemma 4 12B dense, run from bf16 down through Q3_K_M, and the 26B-A4B mixture-of-experts, on my owned DGX Spark as the control and on a rented AWS m8i.12xlarge, Intel Xeon 6 Granite Rapids with Advanced Matrix Extensions (AMX). Total cloud spend: $28.32.
The finding under the finding
The design was a controlled 2x2, one variable per step. Change only the model: the 12B held quality on the Spark at 20 of 39 tasks, and on the honest cross-run metric its 24-task overlap with the baseline reads 13 versus 13 unanimous-pass. Dead even. Change only the substrate: the same 12B moved to the rented Granite Rapids shape and held again at 21 of 39, gaining one task and losing none. Then the addition the data demanded: the 26B mixture-of-experts on the same rented shape tied the 12B at 21 while finishing the identical protocol in 2.24 hours against 5.54. That last one is a configuration comparison, not a single-variable result: architecture, precision, and runtime moved together, with sparse activation the leading explanation. Within the overlapping task set, the larger worker bought no material verified-yield advantage. Its measured advantage was throughput. Correctness lives in the gate.
The number I care most about isn't the yield, it's the false-pass column. A batch system that ships defects quietly is worse than no system, so every validator pass was re-run against held-out tests the model never sees. Zero confirmed false passes, across every configuration measured, including the tiers where quality broke. When the Q3_K_M model got worse, it got worse honestly: more failures routed to the queue, nothing slipping through as false success. The architecture's promise, that the failure mode is a person looks at it, held at every point the lab could measure. That claim extends exactly as far as the validator's detection power. Deterministic doesn't mean complete, and a deterministic gate can reproducibly admit a defect its tests don't detect. The hidden-reference layer is what made the promise falsifiable rather than rhetorical.
So the economics land where the thesis pointed. Per-attempt verified yield ran 51 to 54% across the configurations above the boundary, against the Lab 3 baseline's 47.8% on its own coverage. On the rented shape, the full 117-attempt protocol produced 63 verified bug fixes in 2.24 hours: about nine cents of on-demand compute per verified fix, six cents at the measured one-year Savings Plan rate. And a footnote worth its ink: as measured on July 15, 2026, no classic Reserved Instance offerings existed for the shape at all. The reserved analog is a Savings Plan, which commits dollars, not capacity, so the math here prices unused spend commitment and says so. On owned idle hardware the marginal cost approaches power and operations. Inside an otherwise-unused commitment, the incremental bill can approach zero.
Yes, it fails half the time
The obvious objection: a 54% pass rate is failure half the time, so why not a frontier model or a person? Because the denominator is free and the failures are honest. The 46% that fails costs idle cycles that were already bought and routes to the human who was the fallback before this system existed. The 54% that clears is verified work the human no longer does. The comparison isn't this model versus a better worker. It's this yield versus the zero yield the same committed capacity produced last quarter. And the frontier alternative changes the failure mode, which is the thing this architecture exists to control: a paid-token escalation path fails by accruing charges, while this loop fails by growing a queue. For batch work on committed headroom, a queue is recoverable in a way an open-ended meter is not.
The drift I caught in my own analysis
Midway through, the analysis started treating model determinism as a virtue: celebrating the 12B's zero-flip runs, framing the baseline's verdict flicker at temperature zero as contamination. That inverts the thesis. In Deterministic Code In The Loop (DCITL), the deterministic component is the validator, precisely so the model doesn't have to carry it. A variance-prone worker behind a deterministic gate fails closed against the defects the gate detects, and on idle cycles its retries become mining: every extra pass it flickers into gets checked against held-out evidence before it counts. The metric corruption this drift produces is concrete. The 26B against the baseline on the 29-task overlap reads 17 versus 12 on any-pass, a scary deficit, and 13 versus 12 on unanimous-pass, one task. Any-pass rewards the baseline's own flicker against a zero-flip run. Never compare best-of-n numbers to single-run numbers across different variance profiles. I suspect that quiet corruption infects more published model comparisons than anyone has checked. This rerun also corrected the earlier Loop Control lab's escalation statistics: nine invalid validator checks had booked escalation events for unsolved cases, and corrected, the local model was clearing the solvable work on its own while the frontier tier bought throughput, not correct answers.
The descent found the floor and bracketed it from both sides. Q8_0 and Q4_K_M held the identical pass set as the bf16 original, tier after tier, so the boundary in this environment is a 7.7GB artifact. One rung lower, Q3_K_M broke by five tasks net, triple-confirmed, and the E4B size probe broke from the other axis at 15 of 39, destructively: seven regression verdicts, five of them real broken-worse-than-found edits. Same safety property either way, since the false-pass audit came back zero even at the broken tiers, but what lands in the human queue differs in kind. Intel named a vendor-blessed serving path, on the record, so I tested it and published regardless of direction: real prefill gains on the dense model at 4.7x, slower on bandwidth-bound decode, and unable to execute the mixture-of-experts graph at all. Stock llama.cpp remained the only stack that ran the mixture-of-experts and engaged AMX, confirmed by hardware counters.
What this lab didn't prove matters as much. The boundary doesn't generalize: Q4 as floor and Q3 as break are this environment's coordinates, this model family, this task set, this validator. Two adjacent cells were deliberately left unmeasured and named, because asserting cells you chose not to run is how labs drift into marketing. It ran one task domain, and validators weaker than pytest make the false-pass column harder to trust. The transferable part is the instrument, the descent method, and the audit, and your boundary is an afternoon away. The verdict stands anyway. I watched a 12B model crawl at seven and a half tokens per second and nearly redesigned the lab around speed before the obvious hit: on cycles you already own, slow is a scheduling problem, not a defect. Recovered capacity is real, and it fails honest. A 7.7GB file doing verified maintenance work on hardware that was idling anyway. The sizing specifics are ours; the exercise is yours.
The numbers
Where each layer belongs
| Layer | Placement |
|---|---|
Layer 0 · Compute Compute & Network Fabric The economic object is headroom you already carry, in two distinct forms the page keeps separate: owned idle hardware (the Spark) and unused cloud-spend commitment (the Savings Plan analog; as measured, no classic RI offerings existed for the shape — a Savings Plan commits dollars, not capacity). The lab prices that headroom, not a new purchase. Availability footnote: four AWS launches, zero quota friction, 16 to 29 seconds to SSH. | Retained / Delegated |
Layer 2B · Runtime Application Runtime & Execution The runtime is open source (vLLM, llama.cpp) and the control point is the validator and escalation policy, per the Loop Control ruling. The vendor-blessed OpenVINO backend was measured for scope: 4.7x on prefill, slower on bandwidth-bound decode, and unable to execute the mixture-of-experts graph at all. Stock stacks carried every measured cell. | Retained |
Layer 2C · Reasoning Agentic Infrastructure — The Reasoning Plane Open weights, small and quantized: the boundary worker is a 7.7GB artifact. Nothing in the loop depends on a garden, an API, or a token meter, and the measured cost of that independence on this workload was a few points of yield against the strongest baseline reading. | Retained |
Assessments at the time of the lab
Method and disclosure
Editorial and self-funded: $28.32 of cloud spend, all instances terminated and verified, plus owned hardware (the DGX Spark) for the control runs and the entire quality ladder. Disclosure: Intel is a client of this practice; the lab corresponded with Intel about serving configurations during this program, the exchange is paraphrased where referenced, and no vendor commissioned, funded, or previewed anything. The vendor-blessed serving path was tested because Intel named it, on the record, and the result published regardless of direction.
The instrument is the Lab 3 Loop Control harness, unmodified: 39 tasks (most reconstructed from real bug fixes in httpx, requests, dateutil, more-itertools, arrow, click, and other real libraries, pinned to repo and commit), layered deterministic tests including hidden references, pinned bit-identical across every run by sorted-manifest hash and functional census. Task contents and solutions stay private; the method, source repos, audit design, and every rate, count, and cost ship in the raw detail, including the full chronology with its own mistakes: an instrument-location error in the runbook, a hash recipe that had to be re-anchored, two serving-behavior confounds caught at smoke time, and one cleanup error on a pre-existing cloud volume, disclosed to the account owner the hour it happened.
The essay on this page was rendered from the lab’s frozen structured record by a model writing under this practice’s voice specification, with a deterministic gate checking every figure against the record before publication. The record is the canonical surface: where prose and record disagree, the record rules.
Download the raw lab detail (Markdown)