# Layer2C Labs — Hands-on AI Infrastructure Validation > The Advisor Bench LLC · Updated 2026-08-09 Layer2C Labs runs hands-on technical validation across the 4+1 AI Infrastructure framework. A lab takes a slice of the stack, a single layer, a combination, or the whole model, builds it for real, and renders a verdict on where authority actually sits, scored against the 4+1 Layer AI Infrastructure Model and the Decision Authority Placement Model (DAPM). Editorial labs are self-funded and live on pages here. Sponsored labs live on their own vendor subdomain. Sponsorship funds the work, not the conclusion. How to use this file: the router below has one thin line per lab. Use the three indexes (by vendor, by layer, by theme) to find candidate lab IDs, then open that lab at https://labs.layer2c.com/labs/ (append `.md` or `.json` to the same path for the machine record). The indexes carry plain-language synonyms so a loose question routes without the Layer2C taxonomy. ## Labs (router) — 14 > Field order: ID | title | vendors | layers (4+1 and/or FC codes) | themes | finding | link - L001 | Borrow the vendor’s plumbing, not its judgment | AWS (cloud), NVIDIA (hw) | 0, 1A, 1B, 2B, 2C | managed-abstraction-cost, designed-partition, authority-placement, ai-factory-economics | The managed data plane is cheap and fast; the cost it hides is the chunking judgment it takes from you. Borrow the plumbing, keep your own chunking. | https://labs.layer2c.com/labs/spark-s3vectors - L002 | Own the weights, or the platform owns you | AWS (cloud), NVIDIA (hw) | 0, 2B, 2C | model-ownership, fine-tuning-limits, ai-factory-economics | For a custom model you own, cloud token pricing disappears and becomes a rented floor; against a floor the owned box wins, and the managed path takes the weights. | https://labs.layer2c.com/labs/fine-tune-economics - L003 | The validator determines done, not the loop | model-agnostic; NVIDIA (hw) | 0, 2B, 2C | loop-control, validator-authority, agentic-repair | About 5% of tasks benefit from same-tier repair; the control point is the deterministic validator plus escalation policy, not the loop. | https://labs.layer2c.com/labs/loop-control - L004 | You can’t automate a process you haven’t encoded | Google Cloud (cloud), NVIDIA (hw) | 2A, 2B, 2C | control-plane-authorship, validator-authority, authority-placement | The control plane held where patterns were owned and encoded, and failed where the model became the author of correctness. The question is who may author the patterns and done-criteria. | https://labs.layer2c.com/labs/migration-control-plane - L005 | Authority you reclaim is authority you run | NVIDIA (hw), Google Cloud (ref) | 0, 1A, 1B, 2A, 2B, 2C, FC-0, FC-1, FC-2A, FC-2B, FC-2C, FC-3, FC-4 | cloud-exit, authority-placement, operating-cost-of-ownership | A cloud-free serve path is provable end to end; every layer moved from Ceded to Retained is a decision you now own and a system you now operate. | https://labs.layer2c.com/labs/vctoa-to-spark - L006 | Put the judgment in the constraints, not the weights | NVIDIA (hw), Google Cloud (ref) | 2B, 2C, 3 | judgment-transfer, fine-tuning-limits, prompt-vs-weights, validator-authority | Fine-tuning captured voice and bounded behavior, not judgment; one paragraph of standing positions in the system prompt beat two fine-tuning rounds. | https://labs.layer2c.com/labs/constraints-not-weights - L007 | The CPU exit is a batch lane, not a serving lane | Google Cloud (cloud), Intel (ref) | 0, 2C | cpu-inference, serving-economics, vendor-claim-scrutiny | Across 22 measured configurations, zero met the interactive latency bar; the Xeon lane delivers throughput or interactive latency, not both. | https://labs.layer2c.com/labs/gemma4-xeon-inference - L008 | Renting the chip was the easy part | Google Cloud (cloud) | 0, 2C | accelerator-lock-in, serving-economics, vendor-claim-scrutiny | Renting the TPU is fast; serving a bring-your-own model on it means adopting Google’s stack or quantizing off-box. The friction was the finding. | https://labs.layer2c.com/labs/gemma4-tpu-inference - L009 | Recovered capacity is real, and it fails honest | NVIDIA (hw), Intel (hw), AWS (cloud) | 0, 2B, 2C | recovered-capacity, small-model-viability, validator-authority | A smaller model on idle capacity cleared real verifiable work under the same deterministic validator; in 585 scored attempts no validator pass failed the held-out checks. | https://labs.layer2c.com/labs/xeon-rematch - L010 | Cede the craft, keep the door | model-agnostic | 2B, 2C | validator-authority, authority-placement | Two AI agents shipped a playable game across a file-based seam: the contract was fully delegable, the verdict was not, and you can only cede judgment you have already sourced. | https://labs.layer2c.com/labs/two-agent-seam - L011 | Buy the labor, not the judgment | NVIDIA (hw), Google (model), OpenAI (model), Anthropic (model) | 0, 2B, 2C | loop-control, validator-authority, agentic-repair, authority-placement, ai-factory-economics | One control architecture across 70 repairs. Within that single harness a frontier model added nothing as the loop's controller and lifted recovery from 14 to 20 as the worker. The deterministic validator still determines done, and the tier that matters sits in the worker seat, not the control seat. | https://labs.layer2c.com/labs/labor-not-judgment - L012 | Buy the harness, not the tier | NVIDIA (hw), Google (model), OpenAI (model), Anthropic (model), AWS (cloud), Google Cloud (cloud) | 0, 2B, 2C | loop-control, ai-factory-economics, validator-authority, agentic-repair, recovered-capacity | Open weights on capacity I already carry clear 17 of 22 repairs for nothing, and the gate names the five they miss. Renting more hardware for the rest is slower and dearer than the API. The minimum that finishes the job is a mini model in an agentic harness, at $3.14. Two Opus generations finish the same 22 for four times that. | https://labs.layer2c.com/labs/harness-or-tier - L013 | The floor is trained, not sized | NVIDIA (hw), Alibaba (model), Mistral AI (model), Zhipu AI (model), Google (model), OpenAI (ref) | 0, 2B, 2C | small-model-viability, model-ownership, ai-factory-economics, agentic-repair, validator-authority, recovered-capacity | A 30B coder model on one box clears 18 of 22 gate-verified repairs. A 31B general model cleared 11. Seven parameters apart, same hardware, same gate, and what separates them is training for tools rather than size. The self-hosted floor is real and lower than expected. The bill just isn't in dollars: 18 of 22 took 11.4 hours of owned hardware against 40 minutes and $3.14 for all 22 on a hosted mini. | https://labs.layer2c.com/labs/self-host-floor - L014 | The second box works. The playbook doesn’t. | NVIDIA (hw), DeepSeek (model), Google (model), Alibaba (model), Mistral AI (model) | 0, 2C | serving-economics, vendor-claim-scrutiny, model-ownership | Two clustered DGX Sparks ran a 156 GiB mixture-of-experts checkpoint to 22 of 22 on the certified repair pool, matching a hosted frontier model at zero metered cost, after the vendor’s own multi-node playbook deadlocked at four concurrent requests. | https://labs.layer2c.com/labs/second-box ## Index — by vendor > Every lab where the vendor's technology was on the bench, tested or referenced. Synonyms bridge a loose name to the canonical one. - NVIDIA — L001, L002, L003, L004, L005, L006, L009, L011, L012, L013, L014 (synonyms: Nvidia, GPU vendor, DGX, DGX Spark, GB10, Grace Blackwell) - AWS — L001, L002, L009, L012 (synonyms: Amazon, Amazon Web Services, Bedrock, S3, S3 Vectors, Titan) - Google Cloud — L004, L005, L006, L007, L008, L012 (synonyms: Google, GCP, Vertex AI, TPU, C4, Gemini, GKE) - Intel — L007, L009 (synonyms: Xeon, Granite Rapids, C4 CPU lane) - OpenAI — L011, L012, L013 (synonyms: GPT, GPT-5, GPT-5.5, GPT-5 mini, ChatGPT, o-series) - Anthropic — L011, L012 (synonyms: Claude, Claude Opus, Opus, Claude Code, Sonnet) - Google — L011, L012, L013, L014 (synonyms: Gemma, Gemma 4, Gemini, DeepMind, Google DeepMind) - Alibaba — L013, L014 (synonyms: Qwen, Qwen3, Qwen3-Coder, Qwen3.6, Tongyi, Qwen Team) - Mistral AI — L013, L014 (synonyms: Devstral, Mistral, Codestral, Mistral Small) - Zhipu AI — L013 (synonyms: GLM, GLM-4.7, GLM-4.7-Flash, GLM-4.5-Air, Z.ai, ChatGLM) - DeepSeek — L014 (synonyms: DeepSeek, DeepSeek V4, DeepSeek V4 Flash, DeepSeek-V4-Flash-0731, V4 Flash) A vendor absent here has no published hands-on lab yet, which is not a negative result. See layer2c.com for its desk assessment. ## Index — by layer (4+1 model) > Answers "is there a lab touching ?", including when the layer is named loosely. Canonical 4+1 names live here, once. - Layer 0 — Compute & Fabric — L001, L002, L003, L005, L007, L008, L009, L011, L012, L013, L014 (loose phrasings: chips, GPUs, accelerators, interconnect, networking fabric, silicon) - Layer 1A — Data Plane · Storage — L001, L005 (loose phrasings: storage, object store, data lake, where the data lives) - Layer 1B — Data Plane · Retrieval — L001, L005 (loose phrasings: embedding, data prep, chunking, RAG, retrieval, vector store) - Layer 1C — Data Plane · Context — (labs: none yet) (loose phrasings: context assembly, grounding, what the model gets to see) - Layer 2A — Operational · Control — L004, L005 (loose phrasings: orchestration, scheduling, infrastructure control, the control plane) - Layer 2B — Operational · Execution — L001, L002, L003, L004, L005, L006, L009, L010, L011, L012, L013 (loose phrasings: running the agent, tool calls, task execution, the runtime, serving) - Layer 2C — Operational · Reasoning — L001, L002, L003, L004, L005, L006, L007, L008, L009, L010, L011, L012, L013, L014 (loose phrasings: reasoning, inference, does the model actually think, judgment, decision quality, how it decides) - Layer 3 — Application (+1) — L006 (loose phrasings: the app, the product surface, end-user experience, where the buyer meets it) ## Index — by layer (Fourth Cloud model) > The Fourth Cloud operating-model axis (FC-0 → FC-4). Not a remap of the 4+1 layers; FC scores do not map one-to-one onto them. - FC-0 — Substrate — L005 (loose phrasings: does the control plane lifecycle the hardware, firmware and node management, the physical foundation) - FC-1 — Context Fabric — L005 (loose phrasings: is metadata queryable, the data fabric, enterprise data estate, context for placement) - FC-2A — Orchestration — L005 (loose phrasings: one scheduler for all workloads, VMs and containers and AI in one control plane, unified orchestration) - FC-2B — Runtime — L005 (loose phrasings: unified runtime, consistent developer experience, execution across workload types) - FC-2C — Reasoning Plane — L005 (loose phrasings: policy-driven placement, does policy derive placement from live metadata, autonomous placement) - FC-3 — Catalog — L005 (loose phrasings: is the catalog governed, application distribution, publish version discover consume) - FC-4 — Integration Fabric — L005 (loose phrasings: event bus, API management, workflow orchestration, connectors without point-to-point) ## Index — by theme (controlled vocabulary + synonym bridge) > Each theme is a controlled tag with the plain-language questions it answers. - authority-placement — L001, L004, L005, L010, L011 answers: where does decision authority actually sit in the stack? · who decides, the vendor or me? · how should I score authority across the layers? - managed-abstraction-cost — L001 answers: what does the managed/convenient path quietly decide for me? · hidden cost of the vendor abstraction · what do I give up for convenience? - ai-factory-economics — L001, L002, L011, L012, L013 answers: what does it cost to run this? · token pricing vs owning the box · cost per answer vs cost per token · when does the owned box win on cost? - model-ownership — L002, L013, L014 answers: should I own the weights or rent tokens? · custom model economics · what happens when I need a model I own? - fine-tuning-limits — L002, L006 answers: is fine-tuning worth it? · what does fine-tuning actually change? · fine-tuning vs prompting - judgment-transfer — L006 answers: can you train a model to think like an expert? · does fine-tuning transfer expertise or judgment? · will a fine-tuned model reason like the person? - prompt-vs-weights — L006 answers: system prompt or fine-tune? · where should the expert rules live? · constraints vs training data - validator-authority — L003, L004, L006, L009, L010, L011, L012, L013 answers: how do you check agent output? · where's the real control point? · how to catch a wrong answer · deterministic validator vs model judge - loop-control — L003, L011, L012 answers: do agent retry loops help? · should the agent try again itself? · when does looping stop working? - agentic-repair — L003, L011, L012, L013 answers: does the agent fix its own mistakes? · error recovery in agents · escalation vs retry - control-plane-authorship — L004 answers: who authors correctness in an LLM-assisted process? · can a model own the patterns and done-criteria? · what has to be encoded before you automate it? - cloud-exit — L005 answers: can I leave the managed cloud? · repatriation, bringing the workload home · cloud-exit cost and effort - operating-cost-of-ownership — L005 answers: what does it cost to run what I reclaimed? · day-two cost of owning the stack · reclaim vs operate - serving-economics — L007, L008, L014 answers: can this substrate serve interactively? · latency vs throughput · is it a serving lane or a batch lane? - cpu-inference — L007 answers: can a CPU serve LLM inference? · Xeon for inference · GPU-comparable CPU claims - accelerator-lock-in — L008 answers: how hard is it to serve my model on this silicon? · bring-your-own model on TPU · lock-in to the vendor stack - recovered-capacity — L009, L012, L013 answers: can I use idle owned hardware or committed spend for real work? · reclaiming headroom between peaks · put idle capacity to work - small-model-viability — L009, L013 answers: can a smaller model clear real work? · is a small local model good enough? · quantized model on verifiable work - designed-partition — L001 answers: how do I split the stack across vendors cleanly? · designed partition vs naive hybrid · where should the vendor boundary fall? - vendor-claim-scrutiny — L007, L008, L014 answers: does the vendor's performance claim hold up? · testing a marketed benchmark · is the positioning real? ## Governed chat Governed chat (vCTOA): ask the research system at any property; every answer carries typed evidence, versions, and limits. Machine door + policy: https://virtual-cto-api-1006606217023.us-east4.run.app/api/v1/vctoa/registry ## Machine-readable corpus Corpus index (JSON): https://labs.layer2c.com/labs/index.json — enumerates every lab with canonical URLs, dates, layers, verdict scope, and assessment join tuples. Per lab: https://labs.layer2c.com/labs/.json (full record) and https://labs.layer2c.com/labs/.md. RSS feed: https://labs.layer2c.com/feed.xml — one item per lab, newest first. Precedence when views disagree: per-lab JSON, then the lab page, then this file. All are generated from one source per build, so they do not drift. Related corpus (the canon): https://layer2c.com/llms.txt (index) and per-vendor structured JSON at https://layer2c.com/assessment/.json. Each lab's assessment join tuples carry the vendor source_url, so lab evidence joins to canon verdicts without scraping. ## Part of the Layer2C research system Layer2C assesses authority. Fourth Cloud assesses operating-model readiness. Labs validates where authority holds through scoped builds. StackBuilder composes architectures from the assessed evidence. | Property | URL | Role | | --- | --- | --- | | Layer2C (canon) | https://layer2c.com | 4+1 vendor authority assessments, scored with DAPM | | Fourth Cloud | https://cloud.layer2c.com | On-prem control-plane readiness, FC-0 to FC-4 | | StackBuilder | https://stackbuilder.layer2c.com | Composes architectures from the assessed evidence | | System map | https://layer2c.com/system | How the four properties relate |