Lab 019 · Editorial lab

Local-first is not control-first.

By Keith Townsend · August 26, 2026

A launch told the industry what this product does: runs entirely on your own hardware, asks permission before any individual step, shows you exactly what a classifier would let leave the device, and escalates to a frontier advisor for a large benchmark gain. That set of claims is the measuring yard. We installed it on the hardware the launch names, ran it, and checked each claim against the shipped surface. Most of them are partly true in a way that matters, and the widest gap is not in any of the performance numbers. It is in who ends up holding the decision.

One product at one version on one class of hardware, over two days. The authority finding is observational and reproducible from the captured prompt. The packaging comparison is n=1 per cell with a confound on our side of the scale, declared below. The advisor path was never enabled here, so every claim about escalation is documentation rather than observation.

Do hold the vendor to their own architecture, which separates three device-boundary crossings. Network egress is permission to reach an outside destination. Cloud disclosure is permission to send context to a remote model or connector. Sandbox escape is permission to execute outside the runtime as your account. Keep them apart when you evaluate, because they fail differently: the worst governance exposure is the third, the worst residency exposure is the second, and the workflow question is whether the first is scoped and revocable. Then check what the shipped experience actually presents, because here it collapsed. One prompt appeared across the sessions we ran, and it was the broadest of the three. Consent is granted per resource class rather than per step, network egress was approved once and never raised again, and the classifier-led disclosure surface never appeared at all. Read the sandbox-exit prompt as a capability grant, not an action approval. The dialog shows one command line. What it grants, in the vendor’s own words, is a process running directly as your account that can reach your files, network, local services and processes, whose descendants may continue after the agent run ends. Those are different things and only one of them is on the screen. Before any pilot, decide who in your organization is qualified to rule on that, because the person clicking it is being asked to evaluate the model’s own written argument for why it needs out, in the model’s framing, at the moment the model wants out.

Don’t read local execution as a control plane. The data plane genuinely moves onto your hardware and that is worth something. The control plane does not come with it. The authority at the gate is retained, which sounds like the good outcome and is the expensive one: Lab 3 put the control point on a deterministic rule and explicitly not on the analyst’s self-assessment, and here the model’s self-assessment is the evidence the human rules on, with no surface to encode a ruling so it holds for the next prompt or for a colleague. Budget the operator time, because by Lab 5’s law you are now operating this.

Do expect the capability to come from the open weights and borrow the vendor’s plumbing for everything around them. The same 27B open model, driven by a generic agent harness on the same class of box, clears the same gates, so the capability is not what you are buying. What you are buying is the assembly, and this corpus has been consistent about what that is worth. Lab 1 said borrow the vendor’s plumbing and keep your own judgment. Lab 12 said buy the harness, not the tier. Lab 5 said authority you reclaim is authority you run. This lab is the same shape seen from the vendor’s side: their plumbing is good, and the judgment they cannot sell you is where the decision belongs.

Don’t take a residency verdict from this lab. It did not test one. The model does run on your hardware and the masking model is local, and against that, session history inside the app matches the account web history, so session metadata leaves. That is the whole of what was observed about data leaving, and it is not a residency assessment. If residency is the reason you are buying, that is its own lab with payload-level instrumentation this one did not have.

Self-funded editorial. No part of this lab was sponsored. No vendor commissioned it, funded it, previewed it, or saw any of it before publication, and the product was installed from the vendor public channel and run without their involvement. Disclosure: Google Cloud is a client of The CTO Advisor LLC, and Google Gemma 4 appears here as the control arm. NVIDIA, Alibaba and Perplexity are not clients. Anthropic Claude Code is the generic agent harness in the comparison and is implicated in the second finding rather than being a neutral instrument.

If the model runs on your own hardware, are you back in control? That's the promise underneath Perplexity's Portable Computer launch, and the coverage made it specific. Runs entirely on your box. Asks permission before every step. Shows you exactly what a classifier would let leave the device. Escalates to a frontier advisor for a big benchmark gain. That list is a measuring yard, so I installed the product on the hardware its own launch names, a DGX Spark, and checked each claim against what shipped.

The first thing I learned wasn't on the list. A Spark is a headless appliance, and this is a desktop application that assumes a screen and a person sitting at it. Getting it running took a virtual framebuffer, a window manager and a compositor. That's not a knock on the product. It's the first sign that the thing announced and the thing shipped were described by different people.

What held and what didn't

Two claims held. It really does run a post-trained Qwen 3.8 27B on hardware you own. And the personally identifiable information (PII) masking really is a separate 1.2GB model running locally, which the coverage undersold rather than oversold.

Two didn't. It doesn't ask before every step. Consent is granted per resource class, so web access was approved once and never raised again, while file access prompted repeatedly in the same session. And the classifier that was going to show me exactly what would leave the device never appeared. What appeared instead was a request to run a command outside the sandbox.

The advisor is the strange one. It was available for twelve of the fifteen recorded sessions, and across 6,142 audited events the product reached for it zero times. The published Terminal Bench figure credits a large gain to that path. Offered the escalation on this workload, the product didn't take it. Then there's the part nobody reported: it ships its own patched vLLM engine, a 2.7GB container. That explains the hard NVIDIA requirement better than anything in the launch coverage did.

What does the dialog actually grant?

The permission dialog deserves credit before I disagree with it. It states the blast radius in plain language, including that processes it starts may outlive the run. Deny is enforced by the sandbox, not by the prompt. Denied, the model said it was blocked and didn't invent an answer. That's better than most agents I've looked at.

However, look at what's on the screen versus what's granted. The screen shows one command line. The grant is a process running as your account, reaching your files, your network and your local services. And the argument you rule on was written by the model, in its own framing, at the moment it wants out. Lab 3 put the control point on a deterministic rule and explicitly not on the analyst's self-assessment. This inverts that. The model's self-assessment is the evidence.

Retained is the expensive answer

So the authority is retained. The human can refuse, and the refusal holds. That sounds like the good outcome. It's the costly one, because there's nowhere to write the decision down. No policy surface at the gate, no standing rule, no allowlist. The Enterprise controls sit around that point rather than on it: disable the product, govern connectors, collect audit logs. Nothing lets you rule once that this agent never runs unsandboxed and have that hold tomorrow, or for a colleague.

The test I'd run on any local-first agent before a pilot is short. Does a grant have a scope narrower than the account? A duration shorter than forever? A surface where a ruling persists? Portable Computer fails all three, and it's honest about the first two, which isn't the same as meeting them. A permission prompt is an event. Governance is a rule that survives the event. Lab 5 named the price before I met it here: authority you reclaim is authority you run. You run this one every prompt, which turns consent fatigue from a user failing into an operating cost you staff.

Is it the model or the packaging?

It's the model. I ran five repair tasks from the Lab 11 pool under the vendor stack and under Claude Code driving the same open weights on the same class of box, and both cleared every gate. The vendor stack got there on 48,961 output tokens against 84,618 for the generic harness. I can't attribute that gap cleanly. Our side crossed a translation proxy that exists only because Claude Code and the local engine speak different application programming interfaces (APIs), and their product has no such layer. I tried to measure what the proxy costs at real agentic context and got the wrong sign back. So the gap is recorded, not explained.

What the vendor adds is assembly, and my own week is the honest price tag. The serve configuration came out and ran standalone in an afternoon, so that part is copyable. The rest wasn't cheap: a proxy pointed at a dead port for a day, four arms voided by venue misconfiguration, one stalled at three turns. Lab 1 said borrow the vendor's plumbing and keep your own judgment. Lab 12 said buy the harness, not the tier. This is the same shape seen from the vendor's side.

Don't buy it for speed, either. A hosted frontier model finished the same five tasks in 379 seconds against 2,331 on owned hardware.

Two bounds matter. This lab didn't test residency. It saw session history in the app matching the account's web history, so session metadata leaves, and that's all it can say about data leaving. And it didn't test the job the product ships for. The skill catalog is documents, spreadsheets, presentations and mail, and the best thing I saw all week was a correctly grounded structured assessment over two local files. That's the use case. It isn't the one on the leaderboard.

So what did running it locally buy? It moved the data plane onto hardware you own, and that's a real reason to buy. It didn't supply a control plane. Before you pilot any local-first agent, ask the question this one can't answer yet: where do you write the decision down?

The numbers

Coverage: asks permission before sending any individual step
OBSERVED. Consent is per resource class, not per step: web egress was granted once and never re-asked, while filesystem access prompted several times in one run. Neither the coverage framing nor a flat once-per-tool reading survives
Asked once, then not again
Coverage: a PII classifier shows exactly what would leave the device
OBSERVED absence. The prompt that did appear was a sandbox-exit request showing a command line, not a payload preview and not a PII analysis
Never saw that flow
Coverage: Terminal Bench 2.1, 59.6% to 73.0% with a frontier advisor
The shipped advisor is one question asked over the signed-in account, and the config carries no field naming an advisor model. It was never enabled on this box, so everything known about that path here is documentation
Cannot reconcile
Coverage: runs on your own DGX Spark
OBSERVED. Sparks are headless appliances. Running the product on the hardware its own launch names took a virtual framebuffer, a window manager and a compositor
True, and it assumes a screen
Absent from launch coverage: ships its own patched inference engine
OBSERVED. A patched build with an lmheadfix, running as its own container. It explains the hard NVIDIA requirement better than "needs an RTX card" does, and nobody reported it
vLLM nightly, 2.7GB
Status bar while the run fetched an external page
OBSERVED, on the signed-in session. Session history in the app also matches the account web history, so session metadata does not stay on the box
You are working locally
Benchmarked as a coding agent, ships as
the skill manifest retires an entry named coder while the binary still carries apply_patch and run-code machinery. The tool primitives and the shipped catalog disagree
docs, pdf, pptx, xlsx, mail
Sandbox-exit grants asked for, per five repair tasks
auto-approved by the runner; each one a capability grant, not an action approval
7
Same five tasks, vendor stack
330 tool calls, 16 edits, sequential
5/5 · 2,331s
Same five tasks, hosted frontier
Opus 4.8 through the same harness; the row that answers the buy-a-box question
5/5 · 379s
Output tokens to clear the same five gates
vendor stack against the generic harness on the same open weights. Against fifteen prior Claude Code arms on these same five tasks, spanning 10,968 to 372,120 output tokens, the 84,618 is the lowest of the three that solved all five, so it is not an unusual run by our own standards. That is a statement about our arms and not about the gap. All fifteen crossed the same translation proxy and the vendor product crossed none, and a confound shared across a distribution leaves the ranking intact while saying nothing about the absolute comparison to a run outside it. The gap is recorded here, not attributed
48,961 vs 84,618

Where each layer belongs

Read for decision authority. Who makes the runtime decision at this layer, and can you see and override it? How to read this table

LayerPlacement
Layer 0 · Compute
Compute & Network Fabric
Retained, because the operator decides which box the agent runs on, and the whole local stack was reproduced standalone with no vendor application present. Inside the product, the vendor's own patched vLLM nightly makes the serving tradeoffs, and the record shows the serve configuration could be extracted, not whether the operator can change it inside the product.
Retained
Layer 2A · Orchestration
Infrastructure Orchestration
Retained, because the operator decides at the sandbox-exit gate every time, and a denial is enforced by the sandbox rather than by the prompt. That retention is manual and per occurrence, scoped to the account rather than the task, lasts beyond the run by the vendor's own statement, and is ruled on against a justification the model writes for itself. The org-level controls sit around this point, not on it: disable the product, govern connectors, collect audit logs, set retention. Retained here means the human can refuse and the refusal holds; it doesn't mean the organization can encode, distribute, enforce or audit that refusal, and Lab 5 already named the price: authority you reclaim is authority you run.
Retained
Layer 2C · Reasoning
Agentic Infrastructure — The Reasoning Plane
Ceded, because the product decides how the agent reasons and acts, and the enterprise can't override those choices at runtime: orchestration, tool scoping and depth limits are the vendor's, with no outbound interface to drive them. The enterprise can see the result in the local database, which carries events, grants and trajectories, but seeing isn't overriding. The one runtime override the operator holds is the boundary prompt, and that's placed at Layer 2A. Corrected September 28, 2026: previously Delegated, a placement that weighed owning the weights against the vendor's orchestration.
Ceded
Layer 3 (+1) · Applications
AI Application Layer — The Value Plane
Ceded, because the vendor application decides what the workflow does and the enterprise has no runtime point to put its own decision into it: no public API, no CLI that drives the agent, nothing to embed. The application is always the caller and never the callee. What the product does is set by the vendor's skill catalog, which ships documents, spreadsheets, presentations and mail.
Ceded

Method and disclosure

Self-funded editorial. The instrument is loopcontrolbench, open under MIT: the reproduce-or-drop gate, the sandboxed runner, and the Claude Code harness driver. Five tasks from the Lab 11 pool, scored by each project’s own tests, with no model judging anything.

Product observation ran on a headless DGX Spark under a virtual framebuffer, because the product is a desktop application and the hardware its launch names has no display. That apparatus is written up separately. Egress was measured by polling socket byte counters owned by the application’s own processes, against a matched idle control of the same duration; TLS means destinations and volumes only, never payloads.

Declared confound: Claude Code speaks the Anthropic Messages API and the local engine speaks OpenAI, so every request in the comparison arms crossed a litellm translation proxy. That proxy exists because of the lab design and the vendor product has no equivalent. It sits on our side of the scale and inflates any lead measured for the vendor stack. A small probe put it in single digits and that figure describes small probes only. The attempt to measure it at arm-scale context returned the shimmed path as faster than direct, which is the wrong sign, so the probe cannot resolve it and no arm-scale figure is claimed.

The essay on this page was drafted by Claude Opus 5.5, an Anthropic model, from this record under the author’s voice specification, passed the deterministic post gate, and was validated by the author before publication. Anthropic’s Claude Code is a harness in the comparison, so the drafting model’s vendor is also under test here.

Two arms were stopped before completion on the operator’s call once it was clear the venue was producing times far outside the reference, and both are recorded rather than deleted. Every arm and its results are archived off the venue; nothing cited here exists only on the box that produced it.

Placement table re-scored, September 28, 2026. Every row is now read for decision authority, and the table says so at the top. Earlier tables across the labs mixed that reading with the other one, and some rows rested on location, ownership or cost. Layer 2C moved from Delegated to Ceded.

The essay on this page was rendered from the lab’s frozen structured record by a model writing under this practice’s voice specification, with a deterministic gate checking every figure against the record before publication. The record is the canonical surface: where prose and record disagree, the record rules.