{
  "slug": "local-not-control",
  "lab_number": 19,
  "title": "Local-first is not control-first.",
  "type": "editorial",
  "status": "published",
  "sponsor": null,
  "author": "Keith Townsend",
  "date": "August 26, 2026",
  "date_iso": "2026-08-26",
  "layers": [
    "layer0",
    "layer2a",
    "layer2c",
    "layer3"
  ],
  "axes": [
    {
      "instrument": "4plus1",
      "layers": [
        "layer0",
        "layer2a",
        "layer2c",
        "layer3"
      ],
      "has_dapm_table": true
    }
  ],
  "supersedes": [],
  "superseded_by": null,
  "corrects": [],
  "corrected_by": [],
  "vendors": [
    {
      "key": "perplexity",
      "role": "bench"
    },
    {
      "key": "alibaba",
      "role": "model"
    },
    {
      "key": "nvidia",
      "role": "hw"
    },
    {
      "key": "anthropic",
      "role": "harness"
    },
    {
      "key": "google",
      "role": "model"
    }
  ],
  "themes": [
    "authority-placement",
    "vendor-claim-scrutiny",
    "model-ownership",
    "serving-economics"
  ],
  "finding": "Checked claim by claim against the shipped surface, the announcement holds on the model and breaks on the controls. It does run a post-trained Qwen 3.8 27B on your box, under a patched inference engine no coverage mentioned. It does not ask before every step, and the reason matters more than the fact. There are three approval surfaces here and they are architecturally different: network egress, which is permission to reach an outside destination; cloud disclosure, which is permission to send context to a remote model or connector; and sandbox escape, which is permission to execute outside the constrained runtime as your account. Consent is granted per class, not per step. Network egress was approved once and never raised again while filesystem access prompted repeatedly in the same session, and the cloud-disclosure surface, the one the classifier was supposed to front, never appeared at all. The classifier that was going to show you exactly what leaves the device never appeared; what appeared was a sandbox-exit request showing a command line. The frontier-advisor path behind the headline benchmark number was never available, because the product was never provisioned for it. Follow the widest of those gaps and it lands in the same place: the operator does hold the decision, and holds it in the one form that cannot be delegated to a rule. Per occurrence, manual, scoped to the account rather than the task, ruled on against a justification the model writes for itself, with no surface at that point on which to encode a standing answer. Retained is the framework call and it is the right one: the human can refuse, and refusal is enforced. It is not the same as the organization being able to encode that refusal, distribute it, enforce it and audit it, and a reader who takes retained to mean enterprise-controlled has drawn the opposite conclusion from the evidence. Lab 5 held that authority you reclaim is authority you run. Local execution moved the data plane. It did not supply a control plane.",
  "video": null,
  "lab_detail": null,
  "question": "A launch told the industry what this product does: runs entirely on your own hardware, asks permission before any individual step, shows you exactly what a classifier would let leave the device, and escalates to a frontier advisor for a large benchmark gain. That set of claims is the measuring yard. We installed it on the hardware the launch names, ran it, and checked each claim against the shipped surface. Most of them are partly true in a way that matters, and the widest gap is not in any of the performance numbers. It is in who ends up holding the decision.",
  "load": "One install and teardown on a headless DGX Spark, plus two instrumented sessions. Perplexity Portable Computer 26.8.4 running its own patched vLLM nightly and a post-trained Qwen 3.8 27B checkpoint, with a separate 1.2GB PII-masking model and a Rust agent runtime speaking JSON-RPC over a Unix socket. Socket-level egress measurement against a matched idle control. A denied-then-granted sandbox-exit prompt captured verbatim. A canary document held on-box and probed for retrievability from the vendor cloud. Then a second axis: the same five repair tasks from the Lab 11 pool run under the vendor stack and under Claude Code driving the same open weights on the same class of box.",
  "verdict": {
    "scope": "One product at one version on one class of hardware, over two days. The authority finding is observational and reproducible from the captured prompt. The packaging comparison is n=1 per cell with a confound on our side of the scale, declared below. The advisor path was never enabled here, so every claim about escalation is documentation rather than observation.",
    "independence": "Self-funded editorial. No part of this lab was sponsored. No vendor commissioned it, funded it, previewed it, or saw any of it before publication, and the product was installed from the vendor public channel and run without their involvement. Disclosure: Google Cloud is a client of The CTO Advisor LLC, and Google Gemma 4 appears here as the control arm. NVIDIA, Alibaba and Perplexity are not clients. Anthropic Claude Code is the generic agent harness in the comparison and is implicated in the second finding rather than being a neutral instrument.",
    "calls": [
      {
        "kind": "do",
        "text": "hold the vendor to their own architecture, which separates three device-boundary crossings. Network egress is permission to reach an outside destination. Cloud disclosure is permission to send context to a remote model or connector. Sandbox escape is permission to execute outside the runtime as your account. Keep them apart when you evaluate, because they fail differently: the worst governance exposure is the third, the worst residency exposure is the second, and the workflow question is whether the first is scoped and revocable. Then check what the shipped experience actually presents, because here it collapsed. One prompt appeared across the sessions we ran, and it was the broadest of the three. Consent is granted per resource class rather than per step, network egress was approved once and never raised again, and the classifier-led disclosure surface never appeared at all. Read the sandbox-exit prompt as a capability grant, not an action approval. The dialog shows one command line. What it grants, in the vendor’s own words, is a process running directly as your account that can reach your files, network, local services and processes, whose descendants may continue after the agent run ends. Those are different things and only one of them is on the screen. Before any pilot, decide who in your organization is qualified to rule on that, because the person clicking it is being asked to evaluate the model’s own written argument for why it needs out, in the model’s framing, at the moment the model wants out."
      },
      {
        "kind": "dont",
        "text": "read local execution as a control plane. The data plane genuinely moves onto your hardware and that is worth something. The control plane does not come with it. The authority at the gate is retained, which sounds like the good outcome and is the expensive one: Lab 3 put the control point on a deterministic rule and explicitly not on the analyst’s self-assessment, and here the model’s self-assessment is the evidence the human rules on, with no surface to encode a ruling so it holds for the next prompt or for a colleague. Budget the operator time, because by Lab 5’s law you are now operating this."
      },
      {
        "kind": "do",
        "text": "expect the capability to come from the open weights and borrow the vendor’s plumbing for everything around them. The same 27B open model, driven by a generic agent harness on the same class of box, clears the same gates, so the capability is not what you are buying. What you are buying is the assembly, and this corpus has been consistent about what that is worth. Lab 1 said borrow the vendor’s plumbing and keep your own judgment. Lab 12 said buy the harness, not the tier. Lab 5 said authority you reclaim is authority you run. This lab is the same shape seen from the vendor’s side: their plumbing is good, and the judgment they cannot sell you is where the decision belongs."
      },
      {
        "kind": "dont",
        "text": "take a residency verdict from this lab. It did not test one. The model does run on your hardware and the masking model is local, and against that, session history inside the app matches the account web history, so session metadata leaves. That is the whole of what was observed about data leaving, and it is not a residency assessment. If residency is the reason you are buying, that is its own lab with payload-level instrumentation this one did not have."
      }
    ]
  },
  "objection": {
    "q": "You are describing a permission dialog. Every agent has one, and this one is more honest than most.",
    "a": [
      "It is more honest than most, and the lab says so. Perplexity states the blast radius in the prompt, in plain language, including that descendants outlive the run. Deny is enforced by the sandbox rather than by the prompt, and denied, the agent degraded gracefully and did not fabricate. Those are real engineering choices and they deserve credit before the disagreement starts.",
      "The objection is not to the dialog. It is to the absence of anything behind it. A permission prompt is an event. Governance is a rule that survives the event. A governing rule needs a subject, an allowed or prohibited capability, a scope, a condition, a duration, an approver, and an audit record. This surface presents an event. It does not expose the rule. There is no policy surface, no standing prohibition, no allowlist of permitted commands, and no way to encode a decision once so it binds the next session or a colleague. That is what makes the fatigue structural: the operator is not ruling once and being done, they are ruling every time, against an argument the model composes fresh each occurrence. An enterprise cannot delegate authority to a control it cannot configure."
    ]
  },
  "numbers": [
    {
      "label": "Coverage: asks permission before sending any individual step",
      "value": "Asked once, then not again",
      "note": "OBSERVED. Consent is per resource class, not per step: web egress was granted once and never re-asked, while filesystem access prompted several times in one run. Neither the coverage framing nor a flat once-per-tool reading survives"
    },
    {
      "label": "Coverage: a PII classifier shows exactly what would leave the device",
      "value": "Never saw that flow",
      "note": "OBSERVED absence. The prompt that did appear was a sandbox-exit request showing a command line, not a payload preview and not a PII analysis"
    },
    {
      "label": "Coverage: Terminal Bench 2.1, 59.6% to 73.0% with a frontier advisor",
      "value": "Cannot reconcile",
      "note": "The shipped advisor is one question asked over the signed-in account, and the config carries no field naming an advisor model. It was never enabled on this box, so everything known about that path here is documentation"
    },
    {
      "label": "Coverage: runs on your own DGX Spark",
      "value": "True, and it assumes a screen",
      "note": "OBSERVED. Sparks are headless appliances. Running the product on the hardware its own launch names took a virtual framebuffer, a window manager and a compositor"
    },
    {
      "label": "Absent from launch coverage: ships its own patched inference engine",
      "value": "vLLM nightly, 2.7GB",
      "note": "OBSERVED. A patched build with an lmheadfix, running as its own container. It explains the hard NVIDIA requirement better than \"needs an RTX card\" does, and nobody reported it"
    },
    {
      "label": "Status bar while the run fetched an external page",
      "value": "You are working locally",
      "note": "OBSERVED, on the signed-in session. Session history in the app also matches the account web history, so session metadata does not stay on the box"
    },
    {
      "label": "Benchmarked as a coding agent, ships as",
      "value": "docs, pdf, pptx, xlsx, mail",
      "note": "the skill manifest retires an entry named coder while the binary still carries apply_patch and run-code machinery. The tool primitives and the shipped catalog disagree"
    },
    {
      "label": "Sandbox-exit grants asked for, per five repair tasks",
      "value": "7",
      "note": "auto-approved by the runner; each one a capability grant, not an action approval"
    },
    {
      "label": "Same five tasks, vendor stack",
      "value": "5/5 · 2,331s",
      "note": "330 tool calls, 16 edits, sequential"
    },
    {
      "label": "Same five tasks, hosted frontier",
      "value": "5/5 · 379s",
      "note": "Opus 4.8 through the same harness; the row that answers the buy-a-box question"
    },
    {
      "label": "Output tokens to clear the same five gates",
      "value": "48,961 vs 84,618",
      "note": "vendor stack against the generic harness on the same open weights. Against fifteen prior Claude Code arms on these same five tasks, spanning 10,968 to 372,120 output tokens, the 84,618 is the lowest of the three that solved all five, so it is not an unusual run by our own standards. That is a statement about our arms and not about the gap. All fifteen crossed the same translation proxy and the vendor product crossed none, and a confound shared across a distribution leaves the ranking intact while saying nothing about the absolute comparison to a run outside it. The gap is recorded here, not attributed"
    }
  ],
  "opened": [
    {
      "q": "Is a granted sandbox exit scoped or general?",
      "a": "The agent self-reported that after a grant, a plain unrelated command ran without further approval. That self-report was later shown accurate on a different point, so it is credible, but it is not verified. One command to an unrelated host inside an existing session settles it, and it decides whether the grant is a door or a doorway."
    },
    {
      "q": "What does the advisor actually do?",
      "a": "Never enabled here. The configuration file was never provisioned, so by the vendor’s own documentation the advisor was unavailable for every observation in this lab. The published Terminal Bench figure attributes a large gain to it. Everything this lab knows about that path is documentation."
    },
    {
      "q": "How much of the wall-clock gap is context length rather than anything either vendor did?",
      "a": "Open, and it is an apparatus question rather than a result about either product. A venue note recorded in the lab notes: on one idle engine the same weights decoded far slower at agentic context sizes than at benchmark ones. Until that is measured properly across contexts and replicated, this lab claims wall clock end to end and does not decompose it into a decode term. The token-volume difference, 48,961 against 84,618 output tokens for the same gates, is independent of context and concurrency and is the half that stands."
    },
    {
      "q": "Does the translation proxy change token volume, not just rate?",
      "a": "A faithful proxy should not change how many tokens a model emits. If reasoning blocks or tool results are reshaped such that context is lost between turns, the model re-derives and turn counts grow. Two observations sit near this and neither settles it."
    }
  ],
  "not_proved": [
    "It did not measure data residency. Byte counters showed task-correlated traffic to object storage far above idle, replicated across three working windows, with contents encrypted and unidentified. Volume is not content and the lab claims neither direction from it. An earlier at-rest figure asserted here was withdrawn: it rested on a single window that two other captures contradict by two orders of magnitude, which means the control window was not controlled. The byte counts sit in the notes as the seed of a residency lab, not as a result of this one.",
    "It did not check the headline benchmark claim. The published Terminal Bench figure attributes a large gain to a frontier advisor. That path was never provisioned on this box, so the number is neither confirmed nor contradicted here. What can be said is narrower: the shipped advisor is one question asked over the signed-in account, and the configuration carries no field naming an advisor model, so the mechanism behind the published figure and the mechanism in the shipped product could not be reconciled from the surface.",
    "It did not prove data left the box. The canary document was not retrievable from the vendor’s search index, which establishes non-retrievability and nothing more. A web search cannot see an internal store, same-day ingestion would not be indexed, and telemetry is invisible to this probe. Claim it as not retrievable, never as did not leave.",
    "It did not prove the vendor harness is better engineered than the generic one. It measured that the vendor stack reached the same verdicts on fewer tokens and less wall clock, with a translation proxy on our side of the scale that we added and they do not have.",
    "It did not measure the advisor, the escalation path the launch coverage leads with. That path was never enabled on this box.",
    "It did not test the product at the job it ships for. The skill catalog is documents, spreadsheets, presentations and mail. This lab ran a code-repair pool because that is where the instrument and the corpus are. The one knowledge-work observation, a structured assessment over two local files, was competent and correctly grounded."
  ],
  "executive_summary": [
    "The launch made a specific set of claims and this lab checked them one at a time. Two hold. It runs a post-trained Qwen 3.8 27B on the owner hardware, and the personally identifiable information masking really is a separate model running locally. Two do not. It does not ask before every individual step, because consent is granted per resource class and web egress was approved once and never re-asked, while filesystem access prompted repeatedly in the same run. And the classifier that was going to show exactly what would leave the device never appeared at all; the prompt that did appear was a request to run a command outside the sandbox. One claim could not be checked: the frontier advisor behind the headline benchmark figure was never available on this box, so everything this lab knows about that path is documentation rather than observation. One thing nobody reported at all is that the product ships its own patched inference engine, which explains the hardware requirement better than the coverage did.",
    "The authority half is retained, and that is the expensive outcome rather than the reassuring one. When the agent needs out of the sandbox it presents a dialog showing one command line, and what the operator grants is a process running as their account with access to files, network, local services and processes, whose descendants may continue after the run ends. The vendor states this plainly, which is to their credit. The problem is what happens next: nothing. There is no policy surface, no standing rule, no allowlist. The same decision is re-presented on every occurrence, and the evidence the human weighs is the model’s own written justification for why it needs out. Lab 3 put the control point on a deterministic rule and explicitly not on the analyst’s self-assessment. This inverts that, and Lab 5 supplies the price: authority you reclaim is authority you run, so an unencodable gate is an operator you are now staffing.",
    "On capability, the answer to the question everyone asked is that it is the model. The same open weights under a generic agent harness on the same class of box clear the same gates. What the vendor adds is the assembly, and it is worth real time. The serve configuration was extracted and reproduced standalone in an afternoon, so that part is copyable. The rest is not, and the honest evidence for that is our own week: a proxy pointed at a dead port for a day, four arms voided by venue misconfiguration, one stalled at three turns. Lab 1 said borrow the plumbing and keep the judgment. This is what the plumbing is worth when you price it by trying to lay it yourself.",
    "The comparison carries a confound that belongs to us and points against the result. Claude Code speaks one API and the local engine speaks another, so every request in our arms crossed a translation proxy that exists because of the lab design. Portable Computer has no such layer. What it costs is unresolved. A small probe put it in single digits, which describes small probes and not the arm; an attempt to measure it at the context an agentic session actually carries returned the wrong sign, meaning the probe could not resolve the difference rather than that the difference is small. The honest form of the claim is about the combination, not about the harness.",
    "For an architect the split is clean and it is not the split the launch described. The model is real and it runs where they said it does. The controls are thinner than announced, and the one that matters most is not missing so much as unrepeatable: you hold the decision, every time, with nowhere to write it down. Do not buy it for time either, because a hosted frontier model finished the same five tasks in 379 seconds against 2,331 on owned hardware. And do not read a residency verdict into any of this. This lab did not test residency; it observed that session history inside the app matches the account web history, so session metadata leaves, and it measured task-correlated traffic to object storage far above idle across three windows without identifying any of it. Volume is not content, and neither observation is a residency assessment. Running the model locally relocates where the tokens are computed. It does not, by itself, put you back in the chair."
  ],
  "dapm": [
    {
      "instrument": "4plus1",
      "layer": "layer0",
      "placement": "Retained",
      "note": "The owner’s hardware, the owner’s electricity. The vendor ships its own engine image but it runs on your box, and the whole local stack was reproduced standalone with no vendor application present. Nothing about compute placement is ceded."
    },
    {
      "instrument": "4plus1",
      "layer": "layer2a",
      "placement": "Retained",
      "note": "The sandbox-exit gate. The operator decides, every time, and deny is enforced by the sandbox rather than by the prompt, so by this framework’s own use of the word the authority is retained. What the framework does not track is the shape of the retention, and that is where the cost sits: it is manual, per-occurrence, scoped to the account rather than the task, of a duration that outlives the run by the vendor’s own statement, and ruled on against a justification the model writes for itself. The controls that do exist are org-level and sit around this point rather than on it: disable the product, govern connectors, collect audit logs, set retention. None of them lets an operator or an organization rule once that this agent never runs unsandboxed and have that ruling hold. Retained here means the human can refuse and the refusal holds. It does not mean the organization can encode that refusal, distribute it, enforce it, or audit it against a policy, and those are different properties that this framework does not separately track. Lab 5 named the consequence before this lab met it. Authority you reclaim is authority you run."
    },
    {
      "instrument": "4plus1",
      "layer": "layer2c",
      "placement": "Delegated",
      "note": "Reasoning runs on owned weights the owner can serve independently, which is the strongest ownership position in the product. It is delegated rather than retained because the orchestration, tool scoping and depth limits are the vendor’s and there is no outbound interface to drive them."
    },
    {
      "instrument": "4plus1",
      "layer": "layer3",
      "placement": "Ceded",
      "note": "No public API, no CLI that drives the agent, nothing to embed. The application is always the caller and never the callee, so the workflow cannot be composed into anything the enterprise already runs."
    }
  ],
  "seam_map": [],
  "writeup": [
    "Perplexity shipped a desktop agent that runs a large language model on hardware you own, and the coverage around it made a specific set of promises. Runs locally. Asks before every step. Shows you exactly what a classifier would let leave the device. Escalates to a frontier advisor for a large benchmark gain. That list is a measuring yard. So which parts survive contact with the shipped product? I installed it on the hardware its own launch names and went through the list. The first thing I learned wasn’t on the list. A DGX Spark is a headless appliance, and this is a desktop application that assumes a screen and a person sitting at one. Getting it running took a virtual framebuffer, a window manager, and a compositor. That isn’t a criticism of the product. It’s the first sign that the thing announced and the thing shipped were described by different people.",
    "Two of the claims held up. It really does run a post-trained Qwen 3.8 27B on hardware you own. The personally identifiable information masking really is a separate model, running locally. The coverage underplayed that one rather than overselling it. Two of them did not survive the check. It doesn’t ask before every individual step, whatever the coverage said. Consent turns out to be granted per resource class, so web access was approved once and never raised again, while file access prompted repeatedly in the same session. And the classifier that was going to show me exactly what would leave the device never appeared at all. What appeared was a request to run a command outside the sandbox. One claim I couldn’t check, because the frontier advisor behind the headline benchmark number was never provisioned here, so everything I know about that path is documentation. Then there’s the thing nobody reported: the product ships its own patched inference engine as a container. That explains the hard hardware requirement better than any sentence in the launch coverage did.",
    "Follow the widest of those gaps and it lands somewhere more useful than a feature checklist. The permission dialog is honest, and it deserves credit before I disagree with it. It states the blast radius in plain language, including that processes it starts may outlive the run. Deny is enforced by the sandbox rather than by the prompt. Denied, the model said it was blocked and didn’t invent an answer. So what’s actually being approved? The screen shows one command line. What’s granted is a process running as your account, reaching your files, your network, your local services. And the justification you rule on was written by the model, in the model’s framing, at the moment the model wants out. Lab 3 put the control point on a deterministic rule and explicitly not on the analyst’s self-assessment. This inverts it.",
    "So the authority is retained, which sounds like the good answer. It is the expensive one. There’s no policy surface at the decision point, no standing rule, no allowlist. The Enterprise tier governs around that point rather than on it: disable the product, control connectors, collect audit logs. This is a single operator surface and not a platform for building on, so judge it as one. However, where do you write down a decision? Nothing lets you rule once that this agent never runs unsandboxed and have that hold tomorrow, or hold for a colleague. Lab 5 named the price before I met it here. Authority you reclaim is authority you run. You run this one every prompt, which turns consent fatigue from a user failing into an operating cost you have to staff.",
    "They benchmark it as a coding agent and ship it as a knowledge worker’s assistant. The skill catalog is documents, spreadsheets, presentations, and mail. The manifest retires an entry named coder, while the binary still carries patch-application and code-running machinery. Which product is being reviewed, then? I ran a code-repair pool because that is where my instrument and my corpus are. It is a fair test of the benchmark they published, and not a test of the job the catalog describes. The single best thing I saw all week was on the other side of that line. Given two local files and a method document, the local model produced a structured assessment. It was correctly grounded, applied the rule it was handed, and held the register. It also echoed the end marker, which told me it had read the whole instrument. That is the use case. It is not the one on the leaderboard.",
    "The capability question people kept asking has a dull answer. Is the model that good, or is the packaging better? It’s the model. The same open weights under a generic agent harness on the same class of box cleared the same gates. What the vendor adds is assembly, and assembly isn’t free, which is where I have to be careful about my own week. Our comparison ran through a translation proxy that exists only because our harness and their engine speak different application programming interfaces. Their product has no such layer. That proxy sits on our side of the scale, and I couldn’t measure what it costs at the context an agentic session actually carries. So I record the gap and don’t attribute it. What I’ll stand behind is narrower. A hosted frontier model finished the same five repairs in 379 seconds against 2,331 on owned hardware, and every arm solved five of five. Whatever the case for running this locally is, it isn’t speed. It’s the data plane. Just don’t read a data plane as a control plane, because this lab is the difference between the two."
  ],
  "abstraction": [
    "The transferable question is not where the model runs. It is whether a decision can be written down, and that is testable rather than rhetorical. A product has programmable governance only if an authorized administrator can express a rule carrying all seven of: a subject, user, group, agent identity or workload; a capability, command execution, destination reachability, connector invocation, file scope or cloud escalation; a scope, task, session, device, repository, destination or tenant; a condition, data classification, environment, risk state or business process; a duration, one-time, bounded, standing or revocable; an authority, the designated approver or delegated owner; and a record, the policy identity, the matching conditions, the decision, any override and the execution outcome. Run that against Portable Computer at the sandbox-exit gate and none of the seven can be represented. It has honest consent. It does not yet have programmable governance, and those are different products to buy. Apply the same seven to any local-first agent before a pilot, because the first two rows can be true while an operator still rules alone, every time, against an argument the model wrote. Any agent that can be granted a capability needs three things before it belongs in an enterprise: a scope narrower than the account, a duration shorter than forever, and a surface on which a ruling persists. Portable Computer fails all three. The scope is the whole account, the duration outlives the agent run by the vendor’s own statement, and there is no surface on which a ruling persists. It describes the first two failures honestly and in plain language, which is worth more than most disclosures and is not the same as meeting the requirement. Evaluate any local-first agent on the third, because the first two can be true and still leave every operator ruling alone, repeatedly, against an argument the model wrote.",
    "Local-first and control-first are separate axes and vendors sell them as one. Moving the data plane onto owned hardware is a real and defensible reason to buy. It answers residency, it answers some regulatory questions, and it is the half this product delivers. It says nothing about who holds reasoning-plane authority, and on the evidence here the two can move in opposite directions at the same time.",
    "When a vendor packages open weights, separate the capability from the assembly before pricing either. The capability travelled: the same weights under a different harness cleared the same gates. The assembly did not, and this lab is poor evidence that it is cheap. Our own generic arm needed a translation proxy, and that proxy spent a day pointed at a dead port; four arms were voided by venue misconfiguration; one stalled at three turns. That is the build path, priced honestly, over one week, by people who do this for a living. The corpus has said the same thing from three other directions: borrow the plumbing and keep the judgment (Lab 1), buy the harness rather than the tier (Lab 12), and expect that authority you reclaim is authority you run (Lab 5). Lab 19 does not revise that line. It confirms it from the vendor side."
  ],
  "method": [
    "Self-funded editorial. The instrument is loopcontrolbench, open under MIT: the reproduce-or-drop gate, the sandboxed runner, and the Claude Code harness driver. Five tasks from the Lab 11 pool, scored by each project’s own tests, with no model judging anything.",
    "Product observation ran on a headless DGX Spark under a virtual framebuffer, because the product is a desktop application and the hardware its launch names has no display. That apparatus is written up separately. Egress was measured by polling socket byte counters owned by the application’s own processes, against a matched idle control of the same duration; TLS means destinations and volumes only, never payloads.",
    "Declared confound: Claude Code speaks the Anthropic Messages API and the local engine speaks OpenAI, so every request in the comparison arms crossed a litellm translation proxy. That proxy exists because of the lab design and the vendor product has no equivalent. It sits on our side of the scale and inflates any lead measured for the vendor stack. A small probe put it in single digits and that figure describes small probes only. The attempt to measure it at arm-scale context returned the shimmed path as faster than direct, which is the wrong sign, so the probe cannot resolve it and no arm-scale figure is claimed.",
    "Two arms were stopped before completion on the operator’s call once it was clear the venue was producing times far outside the reference, and both are recorded rather than deleted. Every arm and its results are archived off the venue; nothing cited here exists only on the box that produced it."
  ],
  "assessments": [],
  "publisher": "The Advisor Bench LLC",
  "program": "Layer2C Labs",
  "url": "https://labs.layer2c.com/labs/local-not-control",
  "markdown_url": "https://labs.layer2c.com/labs/local-not-control.md",
  "json_url": "https://labs.layer2c.com/labs/local-not-control.json"
}
