You can’t price the task without pricing the fine-tune.
By Keith Townsend · August 23, 2026
We set out to measure cost per gate-verified task, own hardware against hosted API, and answer whether an expensive owned box earns its price. The measurement kept stalling on a cost we had minimized: the cost of the model itself. You do not get the model for free. You build datasets, run a tune, evaluate it, throw it away, and try again, and the infrastructure the discarded models burned does not vanish because the models did. Pricing a task means pricing that search. This lab ran the search to the end on one domain and came back not with a number but with the procedure that decides whether the number is worth computing at all.
Measured on localized repair with an executable test: one 22-task pool, one open-model family as the tuning subject, one week of checkpoints and prices. The decision procedure is the deliverable; the code-repair verdict is the worked example that produced it. The open edges are a domain that satisfies all five conditions, repository-scale work without localization, the file-geometry wall, and a clean re-measurement of harness edit-landing once the tool-call parser was corrected.
Do settle compliance first, then run the five-condition test. Compliance is an override, not a condition: if the data cannot leave the VPC, or the workload is regulated or air-gapped, you self-host regardless of capability and the conditions below do not apply. Where the hosted frontier is a legal option, the test gates capability fine-tuning: the model has to fail your gate not your budget, the missing capability has to be behavior a corpus cannot inject, a deterministic gate has to exist, the value and volume have to amortize a fixed cost, and you have to price the whole how including the enforcement the tune does not remove. Cost tuning to distill a cheaper model that matches an expensive one is a separate branch that lives in the volume math, not this test.
Don’t treat the fine-tune as an implementation detail inside a cost-per-task number. It is capital. It amortizes across every task it solves, which makes fine-tuning volume economics: catastrophic over one task, invisible over a million of one shape. Any cost model that expenses the training run and forgets the search that preceded it, and the scrap the search burned, is pricing the wrong thing. This search cost roughly $73 and ninety GPU-hours to buy a two-task gain that a hosted model beats outright.
Don’t assume owning the weights buys you a corpus’s worth of knowledge. Retrieval injects what a model knows. Fine-tuning changes how it acts. If a corpus plus a frontier model clears your gate, the task was a RAG problem wearing a fine-tune costume, and most enterprise domains that look proprietary are exactly that. Fine-tuning’s defensible territory is behavior a corpus cannot supply, which is a narrow intersection, not a default.
Do expect the deterministic gate to survive the tune, and budget it. The fine-tune moves behavior most of the way and a hard invariant still leaks the base model’s prior. This practice’s voice work is the proof: even fine-tuned and prompt-guarded, the model reverts to em-dashes, and only deterministic tooling in the loop enforces the pattern. That gate is a recurring cost that never amortizes. Owning the model does not retire the validator. It shifts work onto it.
Do match the apparatus to the model and get the serving contract right before you read a single score. In the loop the consolidation tune reached 14 of 22 against a fair untimed base of 12, a nudge inside noise. On the harness the story was a serving bug: the wrong tool-call parser silently dropped the model’s native emission, and corrected, the untuned base scored 15 where it had shown six to eleven. Loop and harness are different apparatus, so do not read the loop’s 14 against the harness’s 15. Serving correctness is part of the model contract, and a mechanical preflight, envelope, arm-scale prefill, end-to-end tool smoke, provenance stamp, is now the gate every arm passes first.
Self-funded editorial. Disclosure: Google Cloud is a client of The CTO Advisor LLC, and Google models sit at the center of this lab as both tuning subject (Gemma 4) and benched baselines (Gemini 2.5 and 3.6 Flash, Vertex). Anthropic’s Claude Code is the primary harness and is itself implicated in the findings. Hot Aisle, AMD, NVIDIA, OpenAI, Mistral, Alibaba, and DeepSeek are not clients. No vendor commissioned, funded, or previewed any of this.
What does a solved task cost on hardware you own versus a hosted API? That was the whole question. Run one agentic repair pool both ways, count the dollars, and decide whether an expensive owned box earns its price. The measurement kept stalling on a term I'd quietly set to zero: the model itself.
You don't get the model for free. Testing whether an owner's fine-tune beats a foundation model means building the fine-tune. That means a dataset, a training run, an evaluation, a disappointing result, a data change, and another run. The discarded checkpoints are gone, but the infrastructure they burned isn't. So pricing the task means pricing that search, and the lab ran the search to the end on one domain.
The frontier set the bar at $3.14
Code repair was the worked example. A mini-class hosted model solved the entire 22-task pool for $3.14 in forty minutes. That number decided most of what followed, because the foundation model wasn't failing anyone's gate. It was only failing a budget, and it was already cheaper than anything I could build.
Against that floor, the owner tunes ran. Three of them, imitation from a teacher and an isolated-skill precision set, produced no gate-verified improvement. The imitation tune was worse than nothing. It copied the teacher's brevity without the precision that makes brevity work, and it damaged behaviors the base model already had.
The fourth was the one the arithmetic said to build. A self-consolidation low-rank adaptation (LoRA), trained on the model's own gate-passed loop solutions. It ran, and it worked: 14 of 22 in the loop against a fair untimed base of 12. On 22 tasks, a two-task delta sits inside sampling noise. Call it a nudge. And a nudge that still loses outright to a $3.14 hosted run is the cleanest proof I have that this domain belongs to the foundation model, even when ownership does its job.
Didn't you just tune badly?
That's the fair objection: better data or reinforcement learning (RL) closes it. However, look at what an owner can actually reach. Imitation, isolated skill and consolidation are all measured here. What's left is frontier-scale RL post-training, and you can't buy that as a process. You buy it as a product.
The product is already in the open-weights catalog. Devstral and Qwen3-Coder land their edits where Gemma's family doesn't. DeepSeek V4 Flash swept this pool twice on two clustered Sparks I already owned, for nothing but memory and hours. Getting the capability by choosing a model cost nothing. Getting it by training cost me a week to learn I shouldn't have tried.
The humbling part came mid-campaign. The harness wall I'd blamed on the model for a week was substantially a serving bug. The tool-call parser was silently dropping Gemma's native emission. With the correct parser, the untuned base scored 15 of 22 where it had shown six to eleven. Deterministic tooling I controlled had made the model look incapable. The serving contract is part of the model, and I hadn't been treating it that way.
When does fine-tuning pay?
That's the thing worth keeping, because the failure turned into a procedure. It starts with an override. If the data can't leave your virtual private cloud (VPC), or the workload is regulated or air-gapped, you self-host. No amount of frontier capability changes that, and nothing below applies.
Self-hosting also isn't the same decision as fine-tuning. Off-the-shelf open weights can win it with no tune at all, which is exactly what DeepSeek did here. Fine-tuning is the narrowest door, and five conditions have to hold together before you walk through it. The foundation model fails your gate, not just your budget. The missing capability is behavior, not knowledge. A deterministic gate exists. Value per task is high and volume is large enough to amortize a fixed cost. And you've priced the whole how, including the enforcement the tune doesn't remove.
Code repair meets two of the five. Most enterprise domains that look like fine-tuning problems fail on the second condition. Retrieval injects what a model knows. Fine-tuning changes how it acts. If a corpus plus a frontier model clears your gate, you had a retrieval-augmented generation (RAG) problem wearing a fine-tune costume.
The fifth condition is the one the market skips. The tune doesn't retire the validator. The parser was the in-domain proof, and my own voice work is the other: a fine-tuned, prompt-guarded model still reaches for em-dashes until a gate strips them. The fine-tune is capital that amortizes. The gate is a cost you pay on every task, forever. Owning the model shifts work onto the validator. It doesn't remove it.
The bill nobody prices
The search cost roughly $73 metered and about ninety GPU-hours across six days. That number is real, and it's a decoy. The money measures electricity. The six days measure when the answer showed up.
The honest objection is that nobody watches a progress bar, so waiting is free. That's correct, and it misses the point. Slow hardware doesn't burn attended hours. It sits on the critical path. My labs aren't independent: each one produces the next two or three questions, so the run that misses tonight is the experiments that can't start tomorrow. I built a rig to measure cost per solved task, and it told me cost per solved task was never the constraint. Time to solved task was.
The bounds are real. One domain, one open-model family as the tuning subject, no multi-seed runs. The procedure has only been checked against a domain that fails it, so its value on a passing domain is asserted from mechanism, not measured. Lab 002's bounded-generation result, where an owner LoRA did pay, stands untouched.
So own the model when compliance demands it, when open weights already win, or when all five conditions hold and the volume is yours. Rent the frontier otherwise, which is more often than the pitch admits. And the open question is the one this lab can't answer from code repair: which domain actually fails the frontier's gate and passes the other four?
The numbers
Where each layer belongs
Read for portability. Could you take this layer elsewhere without rebuilding? How to read this table
| Layer | Placement |
|---|---|
Layer 0 · Compute Compute & Network Fabric Commodity NVIDIA hardware, so this layer is Retained whoever owns it. Two altitudes, one placement. The owned boxes carried training, serving, loops, and diagnosis at zero marginal dollars, including a two-Spark DDP run once the fabric proved bandwidth-bound rather than latency-bound. The cluster's highest use surfaced late: unified memory holds an open frontier-scale model that sweeps this pool. Capacity, not capability, is what ownership buys, and a host reboot twice was the price of pushing memory past its wall. | Retained |
Layer 2B · Runtime Application Runtime & Execution Retained, because everything at this layer is code the practice wrote and can run anywhere: the gate, the loop, crash-safe results, the venue preflight and the tool-call parser correction, shipped open under MIT in loopcontrolbench. It never regressed, and the whole finding rests on it. Every lesson that lived as code survived the campaign; every lesson that lived as agent memory had to be paid for again. The runbook is the control plane, and agent memory is a cache. | Retained |
Layer 2C · Reasoning Agentic Infrastructure — The Reasoning Plane Split, because the record shows two real paths for the repair-reasoning seat and the lab adopts neither as the answer; its deliverable is the procedure that picks between them. The open-weights path is Retained: DeepSeek V4 Flash swept the pool twice self-hosted and untuned, and the self-consolidation LoRA is an adapter the owner keeps. The hosted path is Delegated: gpt-5.4-mini solved all 22 for $3.14 through the same harness the lab used to swap models in and out of that seat. Corrected September 28, 2026: previously 'Retained via selection, and priced', which isn't a placement value. | Retained/Delegated |
Method and disclosure
Self-funded editorial. The instrument is loopcontrolbench, open under MIT: the reproduce-or-drop gate, the sandboxed runner, both harness drivers, and the additions this campaign contributed, a mechanical venue preflight (envelope, arm-scale prefill, end-to-end tool smoke, provenance stamp), the trajectory taxonomy analyzer, the land-rate gate, the synthetic precision factory, the loop-solution harvester, and a two-Spark DDP training path. The teacher corpus, the trained adapters, and any vendor-assessment consequences stay proprietary. Cross-lab operating discipline accumulated into specs/BENCH-METHOD.md, including the rules this campaign paid for: the tune must earn its place against a clean baseline, serve at the reference envelope with the correct tool-call parser, verify routing by counting requests, cap the loop at two turns, and kill by PID never by pattern.
The essay on this page was drafted by Claude Opus 5.5, an Anthropic model, from this record under the author’s voice specification, passed the deterministic post gate, and was validated by the author before publication. Anthropic’s Claude Code is the primary harness in this lab, so the drafting model’s vendor is also under test here.
Every invalidated arm is archived with its defect named, not deleted: a proxy that routed to the wrong API surface, a context cap below the reference envelope, a ceiling miscalibrated for contention, and a tool-call parser that dropped the model’s native emission. Serving configuration travels with every result via the venue note and a per-task backend probe. The comparisons are kept distinct by design: base versus tuned is the measurement, tuned versus teacher is distillation fidelity, and local versus managed serving is the economics, which this campaign resolved one branch above where the spec opened it, and then generalized into the decision procedure that is the lab’s finding.
Placement note restated, September 28, 2026. The Layer 0 row previously gave ownership or location as its reason. The placement stands; the note now gives the instrument’s reason, because the instrument places Layer 0 by the hardware and the runtime adopted, not by who owns or rents the capacity: commodity x86 and NVIDIA hardware is Retained, and elsewhere an open, standard runtime is Retained while a proprietary runtime that ties the workload to one vendor’s hardware is Ceded.
Placement table re-scored, September 28, 2026. Every row is now read for portability, and the table says so at the top. Earlier tables across the labs mixed that reading with the other one, and some rows rested on location, ownership or cost. Layer 2C moved from Retained via selection, and priced to Retained / Delegated.
The essay on this page was rendered from the lab’s frozen structured record by a model writing under this practice’s voice specification, with a deterministic gate checking every figure against the record before publication. The record is the canonical surface: where prose and record disagree, the record rules.
Download the raw lab detail (Markdown)