Layer2C Labs

What's ready, what's not, where the boundary lives.

Hands-on technical validation across the 4+1 AI Infrastructure framework. A lab takes a slice of the stack, a single layer, a combination, or the whole model, builds it for real, and renders a verdict on where authority actually sits, scored against the 4+1 model and DAPM.

Part of the Layer2C research system · System map
Canon
Layer2C
Assesses where vendor authority sits across the 4+1 stack.
Readiness
Fourth Cloud
Assesses on-prem control-plane readiness, FC-0 to FC-4.
Validation
Layer2C Labs
Validates where authority actually holds, on real hardware.
Application
StackBuilder
Composes architectures from the assessed evidence.

Labs

Lab 017·Editorial·Published

Beyond CUDA: the lock was never the silicon

Lab 002 ruled that the weights are the asset you keep, and every step of that ruling ran on NVIDIA: the box trained, the box served, and the cloud comparison was NVIDIA-backed. So the verdict carried an untested assumption. Does the workflow that produces and serves owned weights actually require CUDA, or is CUDA just where everyone happens to be standing? This lab reran the Lab 002 fine-tune on a rented AMD MI300X, brought the weights home to Apple silicon, and held everything else constant: same base model, same training set, same frozen thirty questions, same recipe down to the learning rate.

Layer 0 · ComputeLayer 2B · RuntimeLayer 2C · Reasoning
Lab 016·Editorial·Published

They can all write it. Not all of them take direction.

The pitch is that picking the right model tier is the decision that determines whether delegated coding work succeeds. This lab walked a model ladder down from an 80B coder to a 2B, on one from-scratch task judged by a deterministic conformance gate, expecting to find the size where capability breaks. It did not find that. Above roughly 12B every rung wrote a structurally complete, plausibly organized server, and the same 26B model on the same task produced three different approaches in three runs, scoring 7 of 7, then 6 of 7, then 7 of 7, in 12, 79 and 21 minutes. The approach predicted the outcome. The size mostly did not.

Layer 2B · RuntimeLayer 2C · Reasoning
Lab 014·Editorial·Published

The second box works. The playbook doesn’t.

Buy a second small AI box and cluster it, and you get a bigger, faster tier for less than a bigger box costs. That is the pitch. Following the vendor’s documented procedure produces a server that answers one request and dies on four.

Layer 0 · ComputeLayer 2C · Reasoning
Lab 013·Editorial·Published

The floor is trained, not sized

Lab 012 bracketed the instrument-literacy bar without locating it. Gemma 4 31B failed it, a hosted mini cleared it, and everything in between was guesswork. If you already own the hardware, that gap is the whole decision. So I ran the descent: the same 22 certified bug-fix repairs, the same deterministic test gate, five open-weight models on one NVIDIA DGX Spark, each inside the same headless Claude Code harness. Which of them clears the bar Gemma missed, and what does self-hosting actually cost once you stop counting dollars?

Layer 0 · ComputeLayer 2B · RuntimeLayer 2C · Reasoning
Lab 012·Editorial·Published

Buy the harness, not the tier

An architect escalating repair work has rungs available: a free local loop, a paid mid-tier model in that same loop, then a frontier model inside an agentic harness. I priced all of them. Same 22 certified bug-fix repairs, same deterministic test gate, six models from a local Gemma to two generations of Opus, run both as a constrained loop and inside Claude Code. The free local loop clears 17. The paid mid-tier loop clears 16 and 17 across two runs, no better. The harness is what closes the rest, and it does that for a mini model as readily as a frontier one. So which rung is actually load-bearing, and what does each one cost?

Layer 0 · ComputeLayer 2B · RuntimeLayer 2C · Reasoning
Lab 011·Editorial·Published

Buy the labor, not the judgment

The pitch is that a local bug-fix agent needs a frontier tier in the loop to escalate to. I rebuilt the chain as originally intended: one loop, one deterministic test gate, one feedback contract, run unchanged across 70 real bug-fix pull requests, swapping only the model in the worker and controller seats and metering every call. A frontier model diagnosing the local worker's failures, as the loop's controller, recovered no net task a free deterministic feedback loop did not. The same frontier model doing the labor, as the worker, cleared work the local model could not. So the intelligence belongs in the worker seat, and the test still decides done.

Layer 0 · ComputeLayer 2B · RuntimeLayer 2C · Reasoning
Lab 010·Editorial·Published

Cede the craft, keep the door

The pitch: a principal who owns only the outcome can compose two spiky AI agents into one shipped artifact across a file-based seam. The loss condition: a judgment the work needs that lives nowhere in the assembly, invisible until it ships.

Layer 2B · RuntimeLayer 2C · Reasoning
Lab 009·Editorial·Published

Recovered capacity is real, and it fails honest

Enterprises commonly carry idle owned hardware or unused cloud-spend commitments above their operating baseline: headroom producing nothing between peaks. This lab asked one question of that idle capacity: can a smaller model clear real, verifiable work on it, judged by the same deterministic validator that judged the big model? Across three models, four quantization tiers, and two substrates, the answer came back yes, with the load-bearing detail attached: in 585 scored attempts, the audit found no validator pass that failed the available held-out checks. The work that cleared was verified as far as the instrument can see. The work that failed went to a human queue. No token meter ran.

Layer 0 · ComputeLayer 2B · RuntimeLayer 2C · Reasoning
Lab 008·Editorial·Published

Renting the chip was the easy part

Google markets the Tensor Processing Unit (TPU) as the price-performance home for Gemma-class inference, and Lab 007 showed the door opens fast: a chip in minutes, a quota bump in minutes, where the NVIDIA lane says no in seconds. This lab set out to serve a mid-size Gemma 4 mixture-of-experts on the lane and fill a latency matrix. It never filled the matrix, because renting the chip turned out to be the easy part. On the silicon you can actually rent self-serve, a bring-your-own model does not fit, and getting it to serve means adopting Google’s stack or quantizing off-box. The performance was already trusted work in earlier labs. The friction was the finding.

Layer 0 · ComputeLayer 2C · Reasoning
Lab 007·Editorial·Published

The CPU exit is a batch lane, not a serving lane

Google promotes the C4 virtual machine for GPU-comparable inference through a customer claim it publishes and features, Intel’s own posts echo the language, and the supporting performance chart compares the new Xeon to the older Xeon. Meanwhile a custom model no garden will host needs compute you control, and the NVIDIA GPU requests this program filed on June 29 were still unusable two weeks later. This lab put a LoRA-tuned Gemma 4 26B mixture-of-experts on the Xeon lane and measured which workload shapes it can actually carry. Across two serving stacks, two prompt shapes, and concurrency 1 through 16, zero of 22 measured configurations met the interactive latency bar. The shape could deliver throughput or interactive latency, not both.

Layer 0 · ComputeLayer 2C · Reasoning
Lab 006·Editorial·Published

Put the judgment in the constraints, not the weights

The pitch is everywhere: fine-tune a local model on a decade of your published judgment and it becomes you. Lab two ruled "own the weights" and built the kill-criterion that goes with it: if base plus retrieval clears the bar, do not fine-tune. This lab is that criterion firing. Two measured training rounds on an owned DGX Spark lost to their own base model, and both lost to a single paragraph of written positions in the system prompt. Throughout, the box means the compute, serving, and training kept below the platform’s abstraction, and the NVIDIA DGX Spark is that box.

Layer 2B · RuntimeLayer 2C · ReasoningLayer 3 (+1) · Applications
Lab 005·Editorial·Published

Authority you reclaim is authority you run

The pitch every cloud-exit deck makes: leave the managed platform and take control back. This lab tests it on one production application. The Virtual CTO Advisor, all-in on a single cloud, migrates to the box (the DGX Spark, retained compute kept below the platform’s abstraction) until the serve path runs with no cloud credentials in the environment. The question is not whether it can run local. It is how much decision authority actually comes home, and what it costs to hold. The one-line loss: every layer you move from Ceded to Retained is a decision you now own and a system you now operate.

Layer 0 · ComputeLayer 1A · StorageLayer 1B · RetrievalLayer 2A · OrchestrationLayer 2B · RuntimeLayer 2C · ReasoningFC-0 · SubstrateFC-1 · Context+5
Lab 004·Editorial·Published

You can’t automate a process you haven’t encoded

I handed a frontier model my migration control-plane operating model and let it build against my own production estate. The control plane did not fail where the patterns were owned and encoded. It failed where the model became the author of correctness. The original question was whether a migration could be metered under the model. The better question the run discovered is who may author the patterns, the validators, and the done criteria in an LLM-assisted control plane.

Layer 2A · OrchestrationLayer 2B · RuntimeLayer 2C · Reasoning
Lab 003·Editorial·Published

The validator determines done, not the loop

The pitch was that a local bug-fix agent needs a frontier tier to escalate to. I built the three-tier chain on a DGX Spark, gated it with a deterministic test harness, and metered every call. Then I audited the harness. Nine of its checks were invalid, and they had booked escalation events that were really unsolved cases sitting on broken tests. Corrected, the credit moves: the local model was clearing the solvable bug fixes on its own, and the frontier tier bought throughput, not correct answers. The test still determines done. That is the part that got more true.

Layer 0 · ComputeLayer 2B · RuntimeLayer 2C · Reasoning
Lab 002·Editorial·Published

Own the weights, or the platform owns you

Lab one found the Spark loses to the cloud on inference. That verdict held only for commodity base models. The moment you need a custom model you own, the cloud stops selling tokens and starts renting you floors, and the managed path takes something you cannot get back: the weights. Throughout, the box means the compute, serving, and training you keep below the platform’s abstraction instead of ceding them, and the NVIDIA DGX Spark is where this lab draws that line.

Layer 0 · ComputeLayer 2B · RuntimeLayer 2C · Reasoning
Lab 001·Editorial·Published

Borrow the vendor’s plumbing, not its judgment

I built a retrieval pipeline across a public-cloud data plane and a local box, the compute I keep below the cloud’s managed abstraction, to map where authority actually sits across the 4+1 stack. The economics were the boring part: eighty-four cents, the cloud faster. The finding worth keeping is what the managed path quietly decides for you, and the two questions the lab now knows to ask.

Layer 0 · ComputeLayer 1A · StorageLayer 1B · RetrievalLayer 2B · RuntimeLayer 2C · Reasoning