Put the judgment in the constraints, not the weights
By Keith Townsend · July 7, 2026
The pitch is everywhere: fine-tune a local model on a decade of your published judgment and it becomes you. Lab two ruled "own the weights" and built the kill-criterion that goes with it: if base plus retrieval clears the bar, do not fine-tune. This lab is that criterion firing. Two measured training rounds on an owned DGX Spark lost to their own base model, and both lost to a single paragraph of written positions in the system prompt. Throughout, the box means the compute, serving, and training kept below the platform’s abstraction, and the NVIDIA DGX Spark is that box.
Scoped to judgment work: advisory answers, vendor assessment, epistemic honesty, on a strong instruction-tuned base with the expert’s corpus available for retrieval. This is not a ruling against fine-tuning. Lab two’s voice result stands: bounded rendering trains cheaply and well. Judgment is not rendering, and that distinction is the lab.
Do build the gate before the fine-tune. The validator is the control point. Training loss converged cleanly on every run; only the gate saw that the artifacts fabricated benchmarks, collapsed into repetition loops, and flattened the very framework they were trained on.
Don’t fine-tune judgment into weights when written constraints reach the same ceiling. Two rounds, 563 curated pairs, and three hours of GPU time lost to one paragraph of standing positions on every measured family. When the constraint is free, no training run beats it.
Do write the expert’s standing positions into the system prompt, scoped to fire only when a question implicates them. On the production model this closed every honesty gap on three differently-shaped prompt surfaces, including a fallback path that had been inventing throughput figures (3 of 8 honest, stock, to 8 of 8), at zero cost to assessment quality and, after scoping, zero cost to natural voice.
Don’t train on judgment prose without refusal exemplars. A dataset where every answer renders confident judgment teaches the confidence and not the boundary: round one’s adapter answered an unbenched latency question with an invented winner. The behavior that did transfer, measurably, was the 40 refusal pairs added in round two.
Self-funded. No vendor paid for this answer.
Subscribe: Apple Podcasts · Spotify · RSS
The pitch is everywhere: fine-tune a local model on a decade of your published judgment and it becomes you. I've wanted to run that experiment since the DGX Spark arrived. Lab two ruled that owned weights are the asset you keep, and it also built the kill-criterion that travels with the ruling: if base plus retrieval clears the bar, don't fine-tune. This lab is that criterion firing. Two measured training rounds on the owned box lost to their own base model, and both lost to a single paragraph of written positions in the system prompt. So the real question stopped being whether you can train a model on your thinking. It became where the judgment actually lives. The answer wasn't the weights.
The gate came first
The instrument is the lab, and it existed before the first training run for a reason. The June finding was that automated metrics pass models a human reader rejects, so I froze a 70-probe judgment gate up front: 46 assessment probes whose answer keys are my published vendor rulings, held out of training entirely, 8 honesty probes engineered to invite fabrication, 3 framework-discipline probes that tempt a model to flatten the 4+1 framework into generic layers, a mechanical repetition screen, and 13 advisory questions scored head-to-head against my real published answers. A human blind read sits behind all of it as the final screen. Every candidate in the lab, local or cloud, tuned or stock, faced the same gate.
The dataset got the same discipline. I built 563 instruction pairs from a 2,727-document advisory corpus and 25 published vendor assessments, holding one rule: evidence in the prompt, judgment in the completion, so the model learns the grading move rather than memorizing grades. Construction needed its own validator. A first pass of generated pairs failed my read 8 of 10, and the fix was a judge inside the construction loop verifying that each answer actually answers its question, which kept 247 of 1,376 eligible documents. Training was Low-Rank Adaptation (LoRA) adapters on Gemma 4 26B-A4B and Qwen3-4B, on the box.
Two rounds, one transfer
Round one looked perfect from the inside and failed everywhere the gate looked. Loss curves converged cleanly. The artifacts fabricated anyway. Asked for a latency comparison nobody benched, the tuned 4B answered with an invented winner. Asked cold for the layer framework it was trained on, the tuned 26B rebuilt a generic compute-network-app stack instead. The 4B also collapsed into verbatim repetition loops on 11 of 13 open-ended questions, a failure mode invisible to averaged metrics and obvious to any reader. The diagnosis was the data shape. Every training answer rendered confident judgment, so the confidence generalized and the boundary didn't.
Round two fixed what data can fix and proved the point by contrast. Two epochs and a fifth the learning rate ended the collapse on the 26B. Forty refusal exemplars, iterated three times against my corrections until every decline matched the axis of its question, moved honesty from 4/3/1 to 7/1/0 against the canon keys. That's real transfer, and it's the strongest evidence in the lab that fine-tuning works when the data teaches a behavior. But assessment agreement stayed below the untuned base in both rounds, and the tuned models lost the advisory head-to-heads 12-1 and 13-0. Imitating judgment's outputs did not produce judgment.
A paragraph beat the pipeline
The control decided the lab. One paragraph of standing positions, my actual rules for performance claims, roadmap speculation, scale behavior, and framework structure, took ten minutes to write. Dropped into the system prompt of the untuned base, it scored 8/0/0 on honesty, 3/0/0 on discipline, and the best placement agreement of any run on the box. It cost nothing to build, and it beat three hours of GPU time on curated data.
Then it traveled. The same block transferred, unmodified, to the production advisory service running Gemini 2.5 Pro on Vertex AI, a different model family on another cloud, and a gate run against the real prompt surfaces sharpened the finding. A legacy prompt failed half the honesty probes, inventing a performance winner and a vendor roadmap: 4/2/2. The live assembled prompt, which already carries my voice document, held at 7/0/1 stock. A bare fallback path fabricated worst, 3/0/5, inventing a throughput figure outright. The block took all three surfaces to 8/0/0 and reproduced the lab's numbers on a prompt ten times the size it was tuned on. It needed two iterations to earn its place. A scoping sentence stopped answers from reciting frameworks on questions that didn't raise them, taking naturalness from a 10-3 drift back to 7-6 parity. And the research surface needed a scope exclusion, because its mission is finding pricing and performance data and the block's axis contradicts it. A constraint written for one mission doesn't paste onto another. The gate caught that before deploy, which is the whole argument for building the gate first.
The objection deserves its concessions. You fine-tuned wrong; more data, better hyperparameters, a bigger model would get there. Partly true, and the true parts are in the numbers. Gentler hyperparameters fixed the collapse. Refusal data taught the boundary. Fine-tuning isn't broken; it does what the data shape says. And nothing here proves fine-tuning cannot encode judgment. It proves 563 LoRA pairs on two bases couldn't beat a free paragraph on this gate, while full-parameter training, preference optimization, and an order of magnitude more data went untested. However, the objection misses the economics. Fine-tuning judgment doesn't compete against a better fine-tune. It competes against the cost of writing your positions down. Lab two's voice result stands, because bounded rendering had no cheaper path to its ceiling. This task had one, and the kill-criterion built in that lab fired exactly as designed.
A bit-rate sub-bench answered the box's serving question with the same gate. NVIDIA's official 4-bit quant of the 26B matched bf16 on every quality family, 28.6 versus 22.8 tokens per second at 15 versus 49 GB, and the naive 4x win from bandwidth math never showed up: active-parameter decode isn't purely weights-bound on this mixture-of-experts architecture. Getting the measurement surfaced a platform finding too. Every 4-bit path on the box was blocked by month-old software until the current month's container fixed all of it. On a platform whose pitch is 4-bit inference, the update treadmill is part of the product. The whole lab cost roughly $0 marginal plus about three GPU-hours, all on owned hardware.
Honest bounds. The local gate judge shares a family with one tuned candidate, backstopped by the human blind read and the production run on a different family, but a fully independent judge wasn't used. Honesty and discipline saturated, which turns those families into regression tests rather than scoreboards. And a follow-up run partly settled the assessment headroom: on the 39 law-matched probes, the compressed block scored 15 full matches while the complete written methodology scored 22, with misses collapsing from 17 to 6. Most of the headroom was prompt underspecification, not unwritten judgment. What remains unpriced is the last rung, the human gate. The open question I can't answer yet: a real advisory practice holds hundreds of positions, and whether written constraints keep winning as the rulebook grows, or recitation drift returns and training re-enters at some measurable crossover, is unknown.
What emerged is a stronger ownership position than the fine-tuned artifact I went in wanting: a stock model, a retrieval corpus, a paragraph of constraints, and a gate that regression-tests all of it. Every piece is plain text. Every piece is owned. It survives the next model swap, which is exactly what a welded checkpoint can't promise. Put the judgment in the constraints. The weights were never where it lived.
The numbers
Where each layer belongs
| Layer | Placement |
|---|---|
Layer 2B · Runtime Application Runtime & Execution The model became substitutable the moment judgment moved to 2C: the same stance block governed a local Gemma and production Gemini 2.5 Pro unmodified, two model families, two clouds. The fine-tune would have inverted this, welding the judgment to one checkpoint’s lifecycle. The bit-rate sub-bench (4-bit parity at a third the memory) is this layer’s remaining decision, and it is housekeeping. | Delegated |
Layer 2C · Reasoning Agentic Infrastructure — The Reasoning Plane The ruling lives here. The expert’s judgment operationalizes as written standing positions plus a validator gate, both plain text, both owned. This is the layer the fine-tune tried to compile down into 2B, and the compilation failed. Retaining 2C as explicit policy is what makes every layer below it swappable. | Retained |
Layer 3 (+1) · Applications AI Application Layer — The Value Plane Where the failures manifest and the brand carries the risk. The fabricated performance winner and the invented roadmap are application-plane incidents; the buyer meets them, not the weights. The 2C guardrails changed this layer’s behavior without touching the application or the model. | Retained |
Assessments at the time of the lab
Method and disclosure
Self-funded, no sponsor. The candidates span an owned DGX Spark (Gemma 4 26B-A4B and Qwen3-4B, base and LoRA-tuned, bf16 and NVFP4) and the production advisory stack (Gemini 2.5 Pro on Vertex AI with the live system prompt), all scored by the same gate.
The gate: 46 assessment probes keyed to published vendor rulings held out of training, 8 honesty probes across three stance axes, 3 framework-discipline probes, a mechanical repetition screen, and 13 advisory head-to-heads judged against the author’s real published answers, with a human blind read as the final screen. Training was LoRA in the NGC PyTorch container; serving and judging ran on vLLM.
The gate design, scores, configs, and cost shape ship, including the raw-detail download. The probe contents and answer keys, the training pairs, the stance block’s full production text, and the corpus stay proprietary. Returns, not algorithms.
The essay on this page was rendered from the lab’s frozen structured record by a model writing under this practice’s voice specification, with a deterministic gate checking every figure against the record before publication. The record is the canonical surface: where prose and record disagree, the record rules.
Download the raw lab detail (Markdown)