What's ready, what's not, where the boundary lives.
Hands-on technical validation across the 4+1 AI Infrastructure framework. A lab takes a slice of the stack, a single layer, a combination, or the whole model, builds it for real, and renders a verdict on where authority actually sits, scored against the 4+1 model and DAPM.
Labs
Recovered capacity is real, and it fails honest
Enterprises commonly carry idle owned hardware or unused cloud-spend commitments above their operating baseline: headroom producing nothing between peaks. This lab asked one question of that idle capacity: can a smaller model clear real, verifiable work on it, judged by the same deterministic validator that judged the big model? Across three models, four quantization tiers, and two substrates, the answer came back yes, with the load-bearing detail attached: in 585 scored attempts, the audit found no validator pass that failed the available held-out checks. The work that cleared was verified as far as the instrument can see. The work that failed went to a human queue. No token meter ran.
Renting the chip was the easy part
Google markets the Tensor Processing Unit (TPU) as the price-performance home for Gemma-class inference, and Lab 007 showed the door opens fast: a chip in minutes, a quota bump in minutes, where the NVIDIA lane says no in seconds. This lab set out to serve a mid-size Gemma 4 mixture-of-experts on the lane and fill a latency matrix. It never filled the matrix, because renting the chip turned out to be the easy part. On the silicon you can actually rent self-serve, a bring-your-own model does not fit, and getting it to serve means adopting Google’s stack or quantizing off-box. The performance was already trusted work in earlier labs. The friction was the finding.
The CPU exit is a batch lane, not a serving lane
Google promotes the C4 virtual machine for GPU-comparable inference through a customer claim it publishes and features, Intel’s own posts echo the language, and the supporting performance chart compares the new Xeon to the older Xeon. Meanwhile a custom model no garden will host needs compute you control, and the NVIDIA GPU requests this program filed on June 29 were still unusable two weeks later. This lab put a LoRA-tuned Gemma 4 26B mixture-of-experts on the Xeon lane and measured which workload shapes it can actually carry. Across two serving stacks, two prompt shapes, and concurrency 1 through 16, zero of 22 measured configurations met the interactive latency bar. The shape could deliver throughput or interactive latency, not both.
Put the judgment in the constraints, not the weights
The pitch is everywhere: fine-tune a local model on a decade of your published judgment and it becomes you. Lab two ruled "own the weights" and built the kill-criterion that goes with it: if base plus retrieval clears the bar, do not fine-tune. This lab is that criterion firing. Two measured training rounds on an owned DGX Spark lost to their own base model, and both lost to a single paragraph of written positions in the system prompt. Throughout, the box means the compute, serving, and training kept below the platform’s abstraction, and the NVIDIA DGX Spark is that box.
Authority you reclaim is authority you run
The pitch every cloud-exit deck makes: leave the managed platform and take control back. This lab tests it on one production application. The Virtual CTO Advisor, all-in on a single cloud, migrates to the box (the DGX Spark, retained compute kept below the platform’s abstraction) until the serve path runs with no cloud credentials in the environment. The question is not whether it can run local. It is how much decision authority actually comes home, and what it costs to hold. The one-line loss: every layer you move from Ceded to Retained is a decision you now own and a system you now operate.
You can’t automate a process you haven’t encoded
I handed a frontier model my migration control-plane operating model and let it build against my own production estate. The control plane did not fail where the patterns were owned and encoded. It failed where the model became the author of correctness. The original question was whether a migration could be metered under the model. The better question the run discovered is who may author the patterns, the validators, and the done criteria in an LLM-assisted control plane.
The validator determines done, not the loop
The pitch was that a local bug-fix agent needs a frontier tier to escalate to. I built the three-tier chain on a DGX Spark, gated it with a deterministic test harness, and metered every call. Then I audited the harness. Nine of its checks were invalid, and they had booked escalation events that were really unsolved cases sitting on broken tests. Corrected, the credit moves: the local model was clearing the solvable bug fixes on its own, and the frontier tier bought throughput, not correct answers. The test still determines done. That is the part that got more true.
Own the weights, or the platform owns you
Lab one found the Spark loses to the cloud on inference. That verdict held only for commodity base models. The moment you need a custom model you own, the cloud stops selling tokens and starts renting you floors, and the managed path takes something you cannot get back: the weights. Throughout, the box means the compute, serving, and training you keep below the platform’s abstraction instead of ceding them, and the NVIDIA DGX Spark is where this lab draws that line.
Borrow the vendor’s plumbing, not its judgment
I built a retrieval pipeline across a public-cloud data plane and a local box, the compute I keep below the cloud’s managed abstraction, to map where authority actually sits across the 4+1 stack. The economics were the boring part: eighty-four cents, the cloud faster. The finding worth keeping is what the managed path quietly decides for you, and the two questions the lab now knows to ask.