{
  "corpus": "Layer2C Labs",
  "description": "Hands-on technical validation across the 4+1 AI Infrastructure framework. Each lab renders a verdict on where decision authority sits, scored against the 4+1 model and DAPM.",
  "publisher": "The CTO Advisor LLC",
  "url": "https://labs.layer2c.com",
  "framework": "4+1 AI Infrastructure Model + Decision Authority Placement Model (DAPM)",
  "updated": "2026-07-26",
  "precedence": [
    "labs/<slug>.json",
    "labs/<slug> (HTML)",
    "labs/index.json",
    "llms.txt"
  ],
  "count": 12,
  "labs": [
    {
      "lab_number": 12,
      "identifier": "LAB-012",
      "slug": "harness-or-tier",
      "title": "Buy the harness, not the tier",
      "type": "editorial",
      "status": "published",
      "date": "July 26, 2026",
      "date_iso": "2026-07-26",
      "layers": [
        "layer0",
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "nvidia",
          "role": "hw"
        },
        {
          "key": "google",
          "role": "model"
        },
        {
          "key": "openai",
          "role": "model"
        },
        {
          "key": "anthropic",
          "role": "model"
        },
        {
          "key": "aws",
          "role": "cloud"
        },
        {
          "key": "gcp",
          "role": "cloud"
        }
      ],
      "themes": [
        "loop-control",
        "ai-factory-economics",
        "validator-authority",
        "agentic-repair",
        "recovered-capacity"
      ],
      "finding": "Open weights on capacity I already carry clear 17 of 22 repairs for nothing, and the gate names the five they miss. Renting more hardware for the rest is slower and dearer than the API. The minimum that finishes the job is a mini model in an agentic harness, at $3.14. Two Opus generations finish the same 22 for four times that.",
      "question": "An architect escalating repair work has rungs available: a free local loop, a paid mid-tier model in that same loop, then a frontier model inside an agentic harness. I priced all of them. Same 22 certified bug-fix repairs, same deterministic test gate, six models from a local Gemma to two generations of Opus, run both as a constrained loop and inside Claude Code. The free local loop clears 17. The paid mid-tier loop clears 16 and 17 across two runs, no better. The harness is what closes the rest, and it does that for a mini model as readily as a frontier one. So which rung is actually load-bearing, and what does each one cost?",
      "verdict_scope": "Measured on localized repair with an executable test, one task pool of 22, one week of vendor pricing. The rulings are about where capability enters an agentic stack and what each rung of an escalation ladder costs per verified unit of output. The open edges are repository-scale debugging without localization, domains without an executable evaluator, and variance bounds beyond two runs per arm.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2C · Reasoning",
          "grade": "Runtime Governance Only — Not a Reasoning Plane",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "AWS AI Infrastructure",
          "layer": "Layer 0 · Compute",
          "grade": "Custom Silicon Full Stack",
          "source_url": "https://layer2c.com/assessment/aws"
        },
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 0 · Compute",
          "grade": "TPU + GPU Full Stack",
          "source_url": "https://layer2c.com/assessment/gcp"
        },
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 2C · Reasoning",
          "grade": "Productized Placement",
          "source_url": "https://layer2c.com/assessment/gcp"
        }
      ],
      "url": "https://labs.layer2c.com/labs/harness-or-tier",
      "markdown_url": "https://labs.layer2c.com/labs/harness-or-tier.md",
      "json_url": "https://labs.layer2c.com/labs/harness-or-tier.json"
    },
    {
      "lab_number": 11,
      "identifier": "LAB-011",
      "slug": "labor-not-judgment",
      "title": "Buy the labor, not the judgment",
      "type": "editorial",
      "status": "published",
      "date": "July 22, 2026",
      "date_iso": "2026-07-22",
      "layers": [
        "layer0",
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "nvidia",
          "role": "hw"
        },
        {
          "key": "google",
          "role": "model"
        },
        {
          "key": "openai",
          "role": "model"
        },
        {
          "key": "anthropic",
          "role": "model"
        }
      ],
      "themes": [
        "loop-control",
        "validator-authority",
        "agentic-repair",
        "authority-placement",
        "ai-factory-economics"
      ],
      "finding": "One control architecture across 70 repairs. Within that single harness a frontier model added nothing as the loop's controller and lifted recovery from 14 to 20 as the worker. The deterministic validator still determines done, and the tier that matters sits in the worker seat, not the control seat.",
      "question": "The pitch is that a local bug-fix agent needs a frontier tier in the loop to escalate to. I rebuilt the chain as originally intended: one loop, one deterministic test gate, one feedback contract, run unchanged across 70 real bug-fix pull requests, swapping only the model in the worker and controller seats and metering every call. A frontier model diagnosing the local worker's failures, as the loop's controller, recovered no net task a free deterministic feedback loop did not. The same frontier model doing the labor, as the worker, cleared work the local model could not. So the intelligence belongs in the worker seat, and the test still decides done.",
      "verdict_scope": "Measured on deterministic coding repair, where a test harness gives an unfalsifiable pass/fail, and after the benchmark supplied file-level localization. The rulings are about where the determines-done authority and the required capability tier sit in an agentic loop. The open edges are autonomous repository debugging, domains without an executable evaluator, and larger replication.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2C · Reasoning",
          "grade": "Runtime Governance Only — Not a Reasoning Plane",
          "source_url": "https://layer2c.com/assessment/nvidia"
        }
      ],
      "url": "https://labs.layer2c.com/labs/labor-not-judgment",
      "markdown_url": "https://labs.layer2c.com/labs/labor-not-judgment.md",
      "json_url": "https://labs.layer2c.com/labs/labor-not-judgment.json"
    },
    {
      "lab_number": 10,
      "identifier": "LAB-010",
      "slug": "two-agent-seam",
      "title": "Cede the craft, keep the door",
      "type": "editorial",
      "status": "published",
      "date": null,
      "date_iso": null,
      "layers": [
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [],
      "themes": [
        "validator-authority",
        "authority-placement"
      ],
      "finding": "Two AI agents shipped a playable game across a file-based seam: the contract was fully delegable, the verdict was not, and you can only cede judgment you have already sourced.",
      "question": "The pitch: a principal who owns only the outcome can compose two spiky AI agents into one shipped artifact across a file-based seam. The loss condition: a judgment the work needs that lives nowhere in the assembly, invisible until it ships.",
      "verdict_scope": "At the scale of one browser game, with two current-generation coding and art agents, and a principal expert in neither craft.",
      "assessments": [],
      "url": "https://labs.layer2c.com/labs/two-agent-seam",
      "markdown_url": "https://labs.layer2c.com/labs/two-agent-seam.md",
      "json_url": "https://labs.layer2c.com/labs/two-agent-seam.json"
    },
    {
      "lab_number": 9,
      "identifier": "LAB-009",
      "slug": "xeon-rematch",
      "title": "Recovered capacity is real, and it fails honest",
      "type": "editorial",
      "status": "published",
      "date": "2026-07-16",
      "date_iso": "2026-07-16",
      "layers": [
        "layer0",
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [
        {
          "lab": "loop-control",
          "what": "The escalation statistics. Nine invalid validator checks had booked escalation events for unsolved cases; corrected, the local model was clearing the solvable work on its own and the frontier tier bought throughput, not correct answers."
        }
      ],
      "corrected_by": [],
      "vendors": [
        {
          "key": "nvidia",
          "role": "hw"
        },
        {
          "key": "intel",
          "role": "hw"
        },
        {
          "key": "aws",
          "role": "cloud"
        }
      ],
      "themes": [
        "recovered-capacity",
        "small-model-viability",
        "validator-authority"
      ],
      "finding": "A smaller model on idle capacity cleared real verifiable work under the same deterministic validator; in 585 scored attempts no validator pass failed the held-out checks.",
      "question": "Enterprises commonly carry idle owned hardware or unused cloud-spend commitments above their operating baseline: headroom producing nothing between peaks. This lab asked one question of that idle capacity: can a smaller model clear real, verifiable work on it, judged by the same deterministic validator that judged the big model? Across three models, four quantization tiers, and two substrates, the answer came back yes, with the load-bearing detail attached: in 585 scored attempts, the audit found no validator pass that failed the available held-out checks. The work that cleared was verified as far as the instrument can see. The work that failed went to a human queue. No token meter ran.",
      "verdict_scope": "Scoped to falsifiable batch work: tasks with a deterministic validator, here real-library bug fixes judged by executable tests. The safety claim extends exactly as far as the validator’s detection power; deterministic does not mean complete, and a deterministic gate can reproducibly admit a defect its tests do not detect. The ruling answers the one question; the boundary numbers are this environment’s coordinates, not universal constants. The size and shape of the models and compute another team needs is their sizing exercise, and this page ships the instrument and method to run it. Two adjacent cells were deliberately not run and are named in the bound: the 26B MoE at Q4 and the E2B size rung. The ruling does not need them, and asserting cells you chose not to measure is how labs drift into marketing.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2B · Runtime",
          "grade": "NVIDIA Authority — Inference + Agent Runtime",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2C · Reasoning",
          "grade": "Runtime Governance Only — Not a Reasoning Plane",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "Intel AI Infrastructure Portfolio",
          "layer": "Layer 0 · Compute",
          "grade": "CPU Strong, Accelerator Present-but-Thin, No Fabric",
          "source_url": "https://layer2c.com/assessment/intel"
        },
        {
          "instrument": "4plus1",
          "vendor": "Intel AI Infrastructure Portfolio",
          "layer": "Layer 2B · Runtime",
          "grade": "Open Runtime Tooling, No Managed Service",
          "source_url": "https://layer2c.com/assessment/intel"
        },
        {
          "instrument": "4plus1",
          "vendor": "AWS AI Infrastructure",
          "layer": "Layer 0 · Compute",
          "grade": "Custom Silicon Full Stack",
          "source_url": "https://layer2c.com/assessment/aws"
        }
      ],
      "url": "https://labs.layer2c.com/labs/xeon-rematch",
      "markdown_url": "https://labs.layer2c.com/labs/xeon-rematch.md",
      "json_url": "https://labs.layer2c.com/labs/xeon-rematch.json"
    },
    {
      "lab_number": 8,
      "identifier": "LAB-008",
      "slug": "gemma4-tpu-inference",
      "title": "Renting the chip was the easy part",
      "type": "editorial",
      "status": "published",
      "date": "2026-07-14",
      "date_iso": "2026-07-14",
      "layers": [
        "layer0",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "gcp",
          "role": "cloud"
        }
      ],
      "themes": [
        "accelerator-lock-in",
        "serving-economics",
        "vendor-claim-scrutiny"
      ],
      "finding": "Renting the TPU is fast; serving a bring-your-own model on it means adopting Google’s stack or quantizing off-box. The friction was the finding.",
      "question": "Google markets the Tensor Processing Unit (TPU) as the price-performance home for Gemma-class inference, and Lab 007 showed the door opens fast: a chip in minutes, a quota bump in minutes, where the NVIDIA lane says no in seconds. This lab set out to serve a mid-size Gemma 4 mixture-of-experts on the lane and fill a latency matrix. It never filled the matrix, because renting the chip turned out to be the easy part. On the silicon you can actually rent self-serve, a bring-your-own model does not fit, and getting it to serve means adopting Google’s stack or quantizing off-box. The performance was already trusted work in earlier labs. The friction was the finding.",
      "verdict_scope": "Scoped to the self-serve, on-demand path on a single GCP project, to the reachable v5e silicon (Trillium v6e capacity was dry on the measured date), and to the open-source vLLM-TPU serving stack. It is a ruling about the level of effort to serve a bring-your-own mid-size mixture-of-experts on the lane, not a claim that Google’s own stack cannot serve Gemma. The served model’s performance is referenced from Lab 006 and Lab 007, not re-measured here. This lab measured the friction, not the tokens per second.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 0 · Compute",
          "grade": "TPU + GPU Full Stack",
          "source_url": "https://layer2c.com/assessment/gcp"
        }
      ],
      "url": "https://labs.layer2c.com/labs/gemma4-tpu-inference",
      "markdown_url": "https://labs.layer2c.com/labs/gemma4-tpu-inference.md",
      "json_url": "https://labs.layer2c.com/labs/gemma4-tpu-inference.json"
    },
    {
      "lab_number": 7,
      "identifier": "LAB-007",
      "slug": "gemma4-xeon-inference",
      "title": "The CPU exit is a batch lane, not a serving lane",
      "type": "editorial",
      "status": "published",
      "date": "July 14, 2026",
      "date_iso": "2026-07-14",
      "layers": [
        "layer0",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "gcp",
          "role": "cloud"
        },
        {
          "key": "intel",
          "role": "ref"
        }
      ],
      "themes": [
        "cpu-inference",
        "serving-economics",
        "vendor-claim-scrutiny"
      ],
      "finding": "Across 22 measured configurations, zero met the interactive latency bar; the Xeon lane delivers throughput or interactive latency, not both.",
      "question": "Google promotes the C4 virtual machine for GPU-comparable inference through a customer claim it publishes and features, Intel’s own posts echo the language, and the supporting performance chart compares the new Xeon to the older Xeon. Meanwhile a custom model no garden will host needs compute you control, and the NVIDIA GPU requests this program filed on June 29 were still unusable two weeks later. This lab put a LoRA-tuned Gemma 4 26B mixture-of-experts on the Xeon lane and measured which workload shapes it can actually carry. Across two serving stacks, two prompt shapes, and concurrency 1 through 16, zero of 22 measured configurations met the interactive latency bar. The shape could deliver throughput or interactive latency, not both.",
      "verdict_scope": "Scoped to a 32-vCPU Granite Rapids shape, which is not the biggest C4 that exists but is the biggest this project could reach through the self-serve quota path, and to this class of model: a mid-size sparse mixture-of-experts served as a merged custom model. The ruling classifies workload shapes. It does not bless or condemn any application; an application owner locates their workload in the taxonomy and reads their row. A post-bench TPU probe tested whether GCP’s third compute lane had a quota or a capacity gate; it measured access only, not inference, and the accelerated-lane findings are otherwise NVIDIA-specific.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 0 · Compute",
          "grade": "TPU + GPU Full Stack",
          "source_url": "https://layer2c.com/assessment/gcp"
        }
      ],
      "url": "https://labs.layer2c.com/labs/gemma4-xeon-inference",
      "markdown_url": "https://labs.layer2c.com/labs/gemma4-xeon-inference.md",
      "json_url": "https://labs.layer2c.com/labs/gemma4-xeon-inference.json"
    },
    {
      "lab_number": 6,
      "identifier": "LAB-006",
      "slug": "constraints-not-weights",
      "title": "Put the judgment in the constraints, not the weights",
      "type": "editorial",
      "status": "published",
      "date": "July 7, 2026",
      "date_iso": "2026-07-07",
      "layers": [
        "layer2b",
        "layer2c",
        "layer3"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "nvidia",
          "role": "hw"
        },
        {
          "key": "gcp",
          "role": "ref"
        }
      ],
      "themes": [
        "judgment-transfer",
        "fine-tuning-limits",
        "prompt-vs-weights",
        "validator-authority"
      ],
      "finding": "Fine-tuning captured voice and bounded behavior, not judgment; one paragraph of standing positions in the system prompt beat two fine-tuning rounds.",
      "question": "The pitch is everywhere: fine-tune a local model on a decade of your published judgment and it becomes you. Lab two ruled \"own the weights\" and built the kill-criterion that goes with it: if base plus retrieval clears the bar, do not fine-tune. This lab is that criterion firing. Two measured training rounds on an owned DGX Spark lost to their own base model, and both lost to a single paragraph of written positions in the system prompt. Throughout, the box means the compute, serving, and training kept below the platform’s abstraction, and the NVIDIA DGX Spark is that box.",
      "verdict_scope": "Scoped to judgment work: advisory answers, vendor assessment, epistemic honesty, on a strong instruction-tuned base with the expert’s corpus available for retrieval. This is not a ruling against fine-tuning. Lab two’s voice result stands: bounded rendering trains cheaply and well. Judgment is not rendering, and that distinction is the lab.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2B · Runtime",
          "grade": "NVIDIA Authority — Inference + Agent Runtime",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 2B · Runtime",
          "grade": "Ceded — Model-Integrated Stack",
          "source_url": "https://layer2c.com/assessment/gcp"
        }
      ],
      "url": "https://labs.layer2c.com/labs/constraints-not-weights",
      "markdown_url": "https://labs.layer2c.com/labs/constraints-not-weights.md",
      "json_url": "https://labs.layer2c.com/labs/constraints-not-weights.json"
    },
    {
      "lab_number": 5,
      "identifier": "LAB-005",
      "slug": "vctoa-to-spark",
      "title": "Authority you reclaim is authority you run",
      "type": "editorial",
      "status": "published",
      "date": "July 4, 2026",
      "date_iso": "2026-07-04",
      "layers": [
        "layer0",
        "layer1a",
        "layer1b",
        "layer2a",
        "layer2b",
        "layer2c",
        "fc0",
        "fc1",
        "fc2a",
        "fc2b",
        "fc2c",
        "fc3",
        "fc4"
      ],
      "instruments": [
        "4plus1",
        "fourthcloud"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "nvidia",
          "role": "hw"
        },
        {
          "key": "gcp",
          "role": "ref"
        }
      ],
      "themes": [
        "cloud-exit",
        "authority-placement",
        "operating-cost-of-ownership"
      ],
      "finding": "A cloud-free serve path is provable end to end; every layer moved from Ceded to Retained is a decision you now own and a system you now operate.",
      "question": "The pitch every cloud-exit deck makes: leave the managed platform and take control back. This lab tests it on one production application. The Virtual CTO Advisor, all-in on a single cloud, migrates to the box (the DGX Spark, retained compute kept below the platform’s abstraction) until the serve path runs with no cloud credentials in the environment. The question is not whether it can run local. It is how much decision authority actually comes home, and what it costs to hold. The one-line loss: every layer you move from Ceded to Retained is a decision you now own and a system you now operate.",
      "verdict_scope": "One production workload, one DGX Spark, one owner. A cloud-free serve path proven end to end by executing probes. This is an authority measurement, not a cost claim and not an on-prem-at-scale claim, scoped to the system that actually ran.",
      "assessments": [
        {
          "instrument": "fourthcloud",
          "vendor": "Bounded Kubernetes (k3s on NVIDIA DGX Spark)",
          "layer": "FC-0 · Substrate",
          "grade": "Ceded / Retained · 1–2/4",
          "source_url": "https://cloud.layer2c.com/assessment/bounded-kubernetes-fourthcloud"
        },
        {
          "instrument": "fourthcloud",
          "vendor": "Bounded Kubernetes (k3s on NVIDIA DGX Spark)",
          "layer": "FC-1 · Context",
          "grade": "Retained · 1–3/4",
          "source_url": "https://cloud.layer2c.com/assessment/bounded-kubernetes-fourthcloud"
        },
        {
          "instrument": "fourthcloud",
          "vendor": "Bounded Kubernetes (k3s on NVIDIA DGX Spark)",
          "layer": "FC-2A · Orchestration",
          "grade": "Retained · 1–3/4",
          "source_url": "https://cloud.layer2c.com/assessment/bounded-kubernetes-fourthcloud"
        },
        {
          "instrument": "fourthcloud",
          "vendor": "Bounded Kubernetes (k3s on NVIDIA DGX Spark)",
          "layer": "FC-2B · Runtime",
          "grade": "Retained · 1–3/4",
          "source_url": "https://cloud.layer2c.com/assessment/bounded-kubernetes-fourthcloud"
        },
        {
          "instrument": "fourthcloud",
          "vendor": "Bounded Kubernetes (k3s on NVIDIA DGX Spark)",
          "layer": "FC-2C · Reasoning",
          "grade": "Retained · 0/4",
          "source_url": "https://cloud.layer2c.com/assessment/bounded-kubernetes-fourthcloud"
        },
        {
          "instrument": "fourthcloud",
          "vendor": "Bounded Kubernetes (k3s on NVIDIA DGX Spark)",
          "layer": "FC-3 · Catalog",
          "grade": "Retained · 1–2/4",
          "source_url": "https://cloud.layer2c.com/assessment/bounded-kubernetes-fourthcloud"
        },
        {
          "instrument": "fourthcloud",
          "vendor": "Bounded Kubernetes (k3s on NVIDIA DGX Spark)",
          "layer": "FC-4 · Integration",
          "grade": "Retained · 1–2/4",
          "source_url": "https://cloud.layer2c.com/assessment/bounded-kubernetes-fourthcloud"
        }
      ],
      "url": "https://labs.layer2c.com/labs/vctoa-to-spark",
      "markdown_url": "https://labs.layer2c.com/labs/vctoa-to-spark.md",
      "json_url": "https://labs.layer2c.com/labs/vctoa-to-spark.json"
    },
    {
      "lab_number": 4,
      "identifier": "LAB-004",
      "slug": "migration-control-plane",
      "title": "You can’t automate a process you haven’t encoded",
      "type": "editorial",
      "status": "published",
      "date": "July 2, 2026",
      "date_iso": "2026-07-02",
      "layers": [
        "layer2a",
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "gcp",
          "role": "cloud"
        },
        {
          "key": "nvidia",
          "role": "hw"
        }
      ],
      "themes": [
        "control-plane-authorship",
        "validator-authority",
        "authority-placement"
      ],
      "finding": "The control plane held where patterns were owned and encoded, and failed where the model became the author of correctness. The question is who may author the patterns and done-criteria.",
      "question": "I handed a frontier model my migration control-plane operating model and let it build against my own production estate. The control plane did not fail where the patterns were owned and encoded. It failed where the model became the author of correctness. The original question was whether a migration could be metered under the model. The better question the run discovered is who may author the patterns, the validators, and the done criteria in an LLM-assisted control plane.",
      "verdict_scope": "One estate, one owner sitting next to the evidence, intake and construction only. No migration ran, and nothing here says the control-plane model fails when humans author the patterns. The finding is about who may author them, and it was earned by watching a frontier model try.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 1A · Storage",
          "grade": "Ceded — Model-Powered Governance",
          "source_url": "https://layer2c.com/assessment/gcp"
        },
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 1B · Retrieval",
          "grade": "Ceded - Model Prep & Managed Retrieval",
          "source_url": "https://layer2c.com/assessment/gcp"
        },
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 2B · Runtime",
          "grade": "Ceded — Model-Integrated Stack",
          "source_url": "https://layer2c.com/assessment/gcp"
        },
        {
          "instrument": "4plus1",
          "vendor": "Google Cloud AI Infrastructure",
          "layer": "Layer 2C · Reasoning",
          "grade": "Ceded — Productized but Captive",
          "source_url": "https://layer2c.com/assessment/gcp"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2C · Reasoning",
          "grade": "Runtime Governance Only — Not a Reasoning Plane",
          "source_url": "https://layer2c.com/assessment/nvidia"
        }
      ],
      "url": "https://labs.layer2c.com/labs/migration-control-plane",
      "markdown_url": "https://labs.layer2c.com/labs/migration-control-plane.md",
      "json_url": "https://labs.layer2c.com/labs/migration-control-plane.json"
    },
    {
      "lab_number": 3,
      "identifier": "LAB-003",
      "slug": "loop-control",
      "title": "The validator determines done, not the loop",
      "type": "editorial",
      "status": "published",
      "date": "July 1, 2026",
      "date_iso": "2026-07-01",
      "layers": [
        "layer0",
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [
        {
          "lab": "xeon-rematch",
          "what": "The escalation statistics. Nine invalid validator checks had booked escalation events for unsolved cases; corrected, the local model was clearing the solvable work on its own and the frontier tier bought throughput, not correct answers."
        }
      ],
      "vendors": [
        {
          "key": "nvidia",
          "role": "hw"
        }
      ],
      "themes": [
        "loop-control",
        "validator-authority",
        "agentic-repair"
      ],
      "finding": "About 5% of tasks benefit from same-tier repair; the control point is the deterministic validator plus escalation policy, not the loop.",
      "question": "The pitch was that a local bug-fix agent needs a frontier tier to escalate to. I built the three-tier chain on a DGX Spark, gated it with a deterministic test harness, and metered every call. Then I audited the harness. Nine of its checks were invalid, and they had booked escalation events that were really unsolved cases sitting on broken tests. Corrected, the credit moves: the local model was clearing the solvable bug fixes on its own, and the frontier tier bought throughput, not correct answers. The test still determines done. That is the part that got more true.",
      "verdict_scope": "Measured on deterministic coding tasks, where a test harness gives an unfalsifiable pass/fail. The rulings are about where the \"determines done\" authority sits in an agentic loop. The open edge is domains without an executable evaluator.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2C · Reasoning",
          "grade": "Runtime Governance Only — Not a Reasoning Plane",
          "source_url": "https://layer2c.com/assessment/nvidia"
        }
      ],
      "url": "https://labs.layer2c.com/labs/loop-control",
      "markdown_url": "https://labs.layer2c.com/labs/loop-control.md",
      "json_url": "https://labs.layer2c.com/labs/loop-control.json"
    },
    {
      "lab_number": 2,
      "identifier": "LAB-002",
      "slug": "fine-tune-economics",
      "title": "Own the weights, or the platform owns you",
      "type": "editorial",
      "status": "published",
      "date": "July 1, 2026",
      "date_iso": "2026-07-01",
      "layers": [
        "layer0",
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "aws",
          "role": "cloud"
        },
        {
          "key": "nvidia",
          "role": "hw"
        }
      ],
      "themes": [
        "model-ownership",
        "fine-tuning-limits",
        "ai-factory-economics"
      ],
      "finding": "For a custom model you own, cloud token pricing disappears and becomes a rented floor; against a floor the owned box wins, and the managed path takes the weights.",
      "question": "Lab one found the Spark loses to the cloud on inference. That verdict held only for commodity base models. The moment you need a custom model you own, the cloud stops selling tokens and starts renting you floors, and the managed path takes something you cannot get back: the weights. Throughout, the box means the compute, serving, and training you keep below the platform’s abstraction instead of ceding them, and the NVIDIA DGX Spark is where this lab draws that line.",
      "verdict_scope": "The economics are arithmetic once you know the floor, and they sit below as evidence. These are the rulings: where authority goes when you customize, and what each managed layer takes in exchange for convenience. One caveat, not a hedge: the box that wins here is a contained, air-cooled, plug-in system. This is not a case for on-prem at scale, where cooling, power, and facilities re-enter the math and this lab did not go.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "AWS AI Infrastructure",
          "layer": "Layer 2B · Runtime",
          "grade": "Delegated / Retained",
          "source_url": "https://layer2c.com/assessment/aws"
        },
        {
          "instrument": "4plus1",
          "vendor": "AWS AI Infrastructure",
          "layer": "Layer 2C · Reasoning",
          "grade": "Intelligence 2C: Delegated | Infra 2C: Implicit",
          "source_url": "https://layer2c.com/assessment/aws"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2B · Runtime",
          "grade": "NVIDIA Authority — Inference + Agent Runtime",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2C · Reasoning",
          "grade": "Runtime Governance Only — Not a Reasoning Plane",
          "source_url": "https://layer2c.com/assessment/nvidia"
        }
      ],
      "url": "https://labs.layer2c.com/labs/fine-tune-economics",
      "markdown_url": "https://labs.layer2c.com/labs/fine-tune-economics.md",
      "json_url": "https://labs.layer2c.com/labs/fine-tune-economics.json"
    },
    {
      "lab_number": 1,
      "identifier": "LAB-001",
      "slug": "spark-s3vectors",
      "title": "Borrow the vendor’s plumbing, not its judgment",
      "type": "editorial",
      "status": "published",
      "date": "June 27, 2026",
      "date_iso": "2026-06-27",
      "layers": [
        "layer0",
        "layer1a",
        "layer1b",
        "layer2b",
        "layer2c"
      ],
      "instruments": [
        "4plus1"
      ],
      "supersedes": [],
      "superseded_by": null,
      "corrects": [],
      "corrected_by": [],
      "vendors": [
        {
          "key": "aws",
          "role": "cloud"
        },
        {
          "key": "nvidia",
          "role": "hw"
        }
      ],
      "themes": [
        "managed-abstraction-cost",
        "designed-partition",
        "authority-placement",
        "ai-factory-economics"
      ],
      "finding": "The managed data plane is cheap and fast; the cost it hides is the chunking judgment it takes from you. Borrow the plumbing, keep your own chunking.",
      "question": "I built a retrieval pipeline across a public-cloud data plane and a local box, the compute I keep below the cloud’s managed abstraction, to map where authority actually sits across the 4+1 stack. The economics were the boring part: eighty-four cents, the cloud faster. The finding worth keeping is what the managed path quietly decides for you, and the two questions the lab now knows to ask.",
      "verdict_scope": "The cost verdict is commodity knowledge, and it sits below as evidence. These are the rulings that change how the stack gets scored: where authority sits, and what you cede without noticing.",
      "assessments": [
        {
          "instrument": "4plus1",
          "vendor": "AWS AI Infrastructure",
          "layer": "Layer 1A · Storage",
          "grade": "Delegated",
          "source_url": "https://layer2c.com/assessment/aws"
        },
        {
          "instrument": "4plus1",
          "vendor": "AWS AI Infrastructure",
          "layer": "Layer 1B · Retrieval",
          "grade": "Delegated",
          "source_url": "https://layer2c.com/assessment/aws"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 0 · Compute",
          "grade": "NVIDIA Strength — Silicon Authority",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2B · Runtime",
          "grade": "NVIDIA Authority — Inference + Agent Runtime",
          "source_url": "https://layer2c.com/assessment/nvidia"
        },
        {
          "instrument": "4plus1",
          "vendor": "NVIDIA AI Platform",
          "layer": "Layer 2C · Reasoning",
          "grade": "Runtime Governance Only — Not a Reasoning Plane",
          "source_url": "https://layer2c.com/assessment/nvidia"
        }
      ],
      "url": "https://labs.layer2c.com/labs/spark-s3vectors",
      "markdown_url": "https://labs.layer2c.com/labs/spark-s3vectors.md",
      "json_url": "https://labs.layer2c.com/labs/spark-s3vectors.json"
    }
  ]
}
