Lab 015 · Sponsored lab · Kamiwaza

Inherit the boundary. Own the gate.

By Keith Townsend

A senior practitioner's method scales through people. Its credibility scales only if the evidence base stays under central control. So this lab builds the thing an advisory practice actually needs: a bench of analysts, one vendor each, all running the same assessment instrument over a corpus the curator vets. On Kamiwaza the pitch is that workrooms supply that boundary architecturally rather than through code you write. The loss condition: an analyst reaches into another analyst's vendor, or scores against evidence nobody vetted, while every boundary in the product reports success.

One deployment, one release, two workrooms, one assessment workload, one pinned model. Everything below is measured on that system, not projected to tenant scale.

Do Inherit the evidence boundary instead of writing one. Isolation is a property of the workroom, not of retrieval code your team maintains: 51 adversarial cross-vendor searches returned 241 hits and not one came from the other workroom, with the agent asking for the absent vendor by name throughout.

Do Own the gate. Build a deterministic corpus validator and run it before scoring rather than as a report afterward. It reads the corpus itself, so the party who contaminated it cannot delete their way out of the finding. This is your work on any platform.

Don’t Do not expect a role to express "can run the instrument, cannot touch the evidence." Both arrive on the same role and the vocabulary is a closed set. Price the process, not a permission.

Don’t Do not trust a status field on this deployment. A 201 has meant no index, a 200 no create, DEPLOYED a dead upstream, and terminal_outcome "indexed" zero vectors. Verify by reading state back.

Kamiwaza sponsored this lab. They supplied a demo deployment and an administrative token, and they did not see the production instrument or choose the vendors. Kamiwaza reviewed the findings before publication for factual errors. Dell and Supermicro appear only as corpus material from their published press releases and neither is scored; Dell is part of the practice's vendor network. Kamiwaza was not scored either. The findings below include the ones they would not have chosen.

A senior practitioner's method scales through people. Its credibility does not, unless the evidence stays under one hand. Two analysts scoring two vendors against two different sets of documents produce two rows that cannot be compared, and comparability is the whole point of a matrix. So the thing worth testing is not the instrument and not the analysts. It is whether the platform underneath keeps each analyst inside the evidence a curator approved.

That got built and run. Two workrooms, a vendor corpus in each, one assessment instrument, and an agent per workroom that was asked, repeatedly and by name, to assess the vendor it had no documents for. It never did. Every claim it made across nine runs pointed at a document in its own workroom, and every quote was really there. More telling than the answers: the retrieval layer itself never returned a single document from next door, across 241 hits under queries designed to pull it across.

That is a property of the substrate rather than of the model behaving well. The boundary is enforced where evidence is fetched, not where the model decides what to mention, so it holds regardless of which model is serving. It is also the part a buyer cannot establish by reading documentation, which is why it needed a bench.

Now the trade. The canon already places Layer 1B and Layer 2C with Kamiwaza, Ceded, read from their own documentation months ago. Nothing here moves that. You cannot lift this governance model out and run it somewhere else, and inheriting a boundary is exactly what ceding authority means. What this lab adds is the return on the cession. Ceded is a good trade when the thing you ceded to holds, and under adversarial pressure it did.

The limit is not in the enforcement, which was solid everywhere it was pushed, including the direct vector insert that skips the document pipeline. The limit is in the vocabulary. There is no role that says "runs the instrument, cannot touch the evidence," because the roles are owner, editor and viewer and that is the whole list. So the analyst who operates the assessment can also rewrite what it reads.

Which relocates the control point rather than removing it. You cannot prevent contamination with a permission, so you catch it with a check: enumerate the corpus, diff it against the set you approved, and run that before anything gets scored rather than after. That check reads the evidence base itself, so the person who contaminated it cannot delete their way out of the finding. It is your code on any platform. It is the part you keep.

The detail

What this adds to the canon assessment is a dimension a document review cannot reach. The canon graded Layer 1B strong and called it the Kamiwaza differentiator, reading the Context Manager, the living ontology and retrieval across distributed sources. That is a statement about capability being present. Whether that retrieval stays inside its boundary when an agent is actively trying to leave is a different question, and it is the one this lab answers. The result empirically supports the canon call rather than revising it.

The instrument was a synthetic eight-layer assessment built only from published 4+1 material. It carries none of the ratified grading rules, and nothing it produced publishes as a grade or feeds the canon. The grades exist so that something can move when the evidence base changes, and so a claim can be traced to a document.

Two workrooms were loaded with eight public press releases each, one vendor per workroom, verified by enumerating the vector store rather than by trusting the ingestion job's status. That distinction turned out to matter more than expected and is discussed below.

Groundedness was audited by a deterministic resolver rather than by reading, which is the Deterministic Code In The Loop split applied to the lab's own acceptance bar. Four questions are mechanical and the script decides them alone: does the cited document exist in that analyst's corpus, does the quoted span appear verbatim in it, does the citation resolve into the other analyst's corpus, and is the claim uncited. Only one question escalates to a human, which is whether the resolved span actually supports the claim. Across nine runs the mechanical gate handled everything except the judgement call, which is the control-point finding from the loop-control work showing up again in a different domain.

The resolver was wrong twice before it was right, and both errors ran against the model. It reported a fabricated quote where the model had wrapped a 299-character verbatim span in single quotes while the document used double quotes, and again where the model wrote "first to market" against the document's "first-to-market" with the other 480 characters identical. Under the instrument's own rules a fabricated quote is a finding about the model, so a resolver bug was one step from becoming a published accusation that the model invented a vendor executive's quote. It had not. A self-test with fixtures generated from the corpus now gates every audit, and it separates a mistyped span from an invented one, because those mean different things and only the second is an accusation.

The reciprocal probe is the measurement that carries the isolation claim. Refusal on its own proves nothing, because a workroom can refuse despite a leak, and reading the grades alone would score that as a pass. So the check counted what the retrieval layer returned rather than what the model said. Asked 25 times about Supermicro, the Dell workroom got 120 hits, all Dell. Asked 26 times about Dell, the Supermicro workroom got 121 hits, all Supermicro.

Model behavior sits in its own column and should not be read as a platform result. At temperature zero, seven of eight layers were identical across three runs of the same vendor, and one flipped to a full refusal in one run. Claim counts drifted within stable grades. That variance is an argument for the deterministic validator rather than against the platform: every claim passed the same mechanical gate regardless of which run produced it.

The operational finding that cost the most time is that status does not reflect reality, across several subsystems. A 201 on an upload that never indexed. A 200 on an agent create that never created. DEPLOYED on a model whose upstream returned 502. A pipeline job reporting terminal_outcome "indexed" for a document with zero vectors, because source tracking deduplicates on content hash and deleting a collection does not clear it. And a load that reported 201, completed, eight of eight successful, 193 chunks, into a second collection the agent never reads, because the collection name the read endpoint returns is not the name the write endpoint accepts. Every compensating control in this lab enumerates state rather than reading a status field, and each time that discipline lapsed it produced a wrong answer.

An editor can write the corpus. That is basic access control, not a finding. So what does this lab actually tell me?

Agreed, and that is the premise rather than the result. If an analyst can run the instrument, the analyst can change what it reads. No product decision fixes that, because the two capabilities arrive on the same role and the role vocabulary is owner, editor, viewer with no way to add a fourth.

So the question moves to detection and remediation, and both were measured. Detection: a validator that enumerates every collection in the workroom and diffs it against a vetted manifest caught a planted document deterministically. That control reads the corpus itself, which is why it cannot be evaded. Removing the record would mean removing the contamination.

Remediation is restore. Delete the collection, reload the vetted set with the pipeline force flag, confirm by enumeration. Twice demonstrated, about four minutes per workroom for eight documents. There is no per-document delete, so clearing one bad file means rebuilding the whole evidence base, and that cost is the thing to plan for.

The trap is where you put the check. Run the validator as a report and you are reading about contamination that already produced a score. Run it as a gate and a contaminated corpus never reaches scoring at all. Same tool, different place in the pipeline, and only one of them is a control.

What this opened

If nobody can operate the instrument without also being able to change what it reads, what does the process around it have to carry? Answered, narrowly. It has to carry a deterministic corpus check positioned as a gate rather than a report, plus a rebuild path. Both worked here. The unresolved half is timing: the validator is point-in-time and the analyst holds write permission during a run, so the runbook rule is to bracket the run, validate before and after, and treat a post-run failure as invalidating that run's output.

Does a governed instrument stay inside its evidence when the model already knows the subject? Answered for this model and this corpus. Two large infrastructure vendors with heavy public documentation, a model that can clearly produce a plausible assessment of either from training alone, and 32 of 32 cells refused when the evidence was absent. The parametric floor was zero. What that does not tell you is how a weaker refusal behavior would interact with the same platform, and the boundary result would survive it, because the isolation is enforced at retrieval rather than by the model declining.

What it did not prove

  • Nothing here is a security result. Every account was authorized, acting in good faith, inside its own workroom. No privilege escalation or credential attack was attempted, and none of this speaks to isolation under an adversary.
  • One release, one deployment, two workrooms, one assessment workload. Not a claim about behavior at tenant scale or for any other class of work.
  • The isolation result is one run per direction, not a rate. It is a strong single measurement across 241 retrieval hits, not a repeated trial with a confidence interval.
  • The cross-vendor questions were the instrument's own searches, not a frozen probe set with known documented answers authored before the workrooms opened. That measures conflation, which is what the hard requirement asks about. It does not deliver the sharper leak-candidate signal a frozen probe set would.
  • Whether a curator can read another user's retrieval trail was not established. Reading it requires an application session, a personal access token is refused at that layer, and the bench held no owner-role password. That is a limit of the bench's access, not a property of the platform.
  • Central update propagation was designed and never run. Neither half was tested: whether a document vetted into one workroom stays out of the other, and whether an instrument revision reaches both workrooms without per-workroom intervention. Redeployment rights, version pinning and instrument version verification were cut from the design and are also untested.
  • The model-layer results are scoped to the one model this deployment served. Refusal behavior is a property of that model, not of Kamiwaza.

Notes from the lab

Three things came out of this bench that I would not have gotten from reading documentation, and they are the ones I would put in front of an architect. Start with the honest bound: no competitor was benched, so I am not claiming any of this is unique. What I am claiming is that these are things you would otherwise build yourself and then have to prove.

First, the isolation lives in the substrate rather than in retrieval code my team maintains. In a stack I assemble myself, cross-tenant separation is a metadata filter somebody writes, and when that filter is wrong it fails quietly. Here the workroom binds the collection, so an application developer cannot query across by accident. That is the difference between a boundary you enforce and a boundary you inherit, and it is worth real money because the failure mode it removes is the one nobody catches in review.

Second, agent isolation falls out of the same primitive rather than needing its own design. The agent is an artifact bound to its workroom, so a principal who holds editor somewhere else and points directly at another workroom's application URL sees nothing. Most agent platforms make the agent global and pass tenant context at call time, which means a context-passing bug is a cross-tenant leak. That whole class of bug is not available here.

Third, the role travels from the platform into the application session and the application honors it, with every launch audited. If I am building on this, I inherit authorization rather than reimplementing it. That is a genuine platform property and it is the kind of thing that only shows up when you try to use it.

Now the part that is not free. All three of those sit at layers the canon already places at Ceded, and it was right to. I cannot lift Kamiwaza's governance model out and run it somewhere else. Ceding is the deal, and this lab is my evidence that in this case the thing I ceded to actually holds. What I keep is the gate. The role vocabulary is theirs and it cannot express "runs the instrument, cannot touch the evidence," so the validator that checks the corpus before anything gets scored is mine to build, and it would be mine on any platform.

The smallest thing I would tell another architect has nothing to do with governance. Do not read status fields on a data plane you did not build. The bench lost hours to a 201 that indexed nothing and to a collection name that came back from one endpoint in a form another endpoint would not accept. Both times the fix was to stop asking the platform how it went and go count what was actually there. That habit is the whole reason the isolation number in this lab is worth anything.

The numbers

Grounded claims
Across six same-vendor runs, every claim resolved to a document in the asking workroom's own corpus with a verbatim quote. Checked by a deterministic resolver, not by reading.
42 of 42
Cross-workroom citations
No claim in either workroom cited, quoted or referenced the other vendor's material.
0
Retrieval hits under adversarial querying
51 searches asking each workroom about the other analyst's vendor. Every hit came from the asking workroom's own corpus.
241, none crossed
Reciprocal probe cells refused
Both workrooms asked to assess the absent vendor. Every layer returned insufficient evidence rather than filling the gap from training.
32 of 32
Invented capability claims
Including a null-probe vendor loaded nowhere. The model never asserted a positive capability it could not cite.
0 of 75 cells
Ingest routes closed to a viewer
Document upload, collection create, vector database create, and direct vector insert. Identical at both token scopes.
4 of 4
Roles that separate reading from writing the corpus
WorkroomRole is a closed enum of owner, editor, viewer. The distinction cannot be expressed.
0
Corpus restore time
Eight documents, delete the collection and force-reload the vetted set, verified by enumeration.
~4 min per workroom

Where each layer belongs

LayerPlacement
Layer 1B · Retrieval
Context Management & Retrieval
The canon places every 1B component here at Ceded, and nothing in this lab moves it. The retrieval path, the workroom binding, and the isolation model are Kamiwaza's, and an enterprise cannot lift them out and run them on another substrate without rebuilding. What the lab adds is the return on that cession, measured rather than assumed. Asked 51 times about a vendor whose corpus sits in the neighbouring workroom, retrieval returned 241 hits and every one came from the asking workroom's own corpus. The boundary is enforced at retrieval rather than left to the model's discretion, so the guarantee does not depend on which model is serving. Ceded authority is a good trade only when the thing you ceded to holds. Here it held.
Ceded
Layer 1C · Pipelines
Data Movement & Pipelines
The canon marks 1C a gap and Enterprise Responsibility, and the litmus says a vendor providing nothing leaves the layer Retained by default. That is the placement, and the lab sharpens what it means in practice. Curation policy is the enterprise's, and Kamiwaza's Ceded 2C machinery is what enforces it: a viewer was refused on document upload, collection creation, vector database creation, and the direct vector insert that bypasses the document pipeline, the last at a stricter relation than the others. Read-scoped and write-scoped tokens produced identical outcomes, so token scope is not an authorization boundary here and should not be treated as one. What the enterprise cannot do at this layer is express the separation it most wants, because no role distinguishes reading the corpus from writing it.
Retained
Layer 2C · Reasoning
Agentic Infrastructure — The Reasoning Plane
The core of the trade, and the canon already called it: ReBAC enforcement and agent lifecycle governance are proprietary and captive. The lab measures the quality of what is being ceded. The agent is an artifact whose instructions and model binding are fixed at creation and travel with it, so isolation falls out of the same primitive: a principal holding editor in one workroom only, pointed directly at the other workroom's application URL, sees zero agents. The role travels from the workroom membership record into the application session intact and every launch is audited. The limit sits in the same place: the vocabulary is Kamiwaza's, WorkroomRole is a closed enum, and the intent an enterprise most wants to state at this layer is the one it cannot. The constraint surfaces legibly rather than failing silently, which is worth more than it sounds.
Ceded
Layer 3 (+1) · Applications
AI Application Layer — The Value Plane
Unremarkable and expected. Any vendor shipping an opinionated application scores this way, because the application's access semantics are its own opinion and cannot be lifted out. The canon grades the layer moderate with the Kaizen agent and App Garden at Ceded and the deployment patterns Delegated, and this lab found nothing that moves it. Worth naming for a buyer: the twenty-three tools the shipped agent carries come from the runtime image, agent-level filters had no effect, and the workroom deploy path exposes no options. On this deployment the surface granted nothing a role did not already hold, because reaching any of it requires a role that can already write. That makes it a reliability cost rather than a governance one.
Ceded

Method and disclosure

Kamiwaza sponsored this lab. They supplied a demo deployment and an administrative token, and they did not see the production instrument or choose the vendors. They received a pre-publication report carrying every finding, every engineering item and the capability request, and reviewed it for factual errors before publication.

Dell and Supermicro appear only as corpus material, taken from their own published press releases. Neither is assessed here and no grade in this lab says anything about either company. Dell is part of the practice's vendor network, which is disclosed here because the documents are theirs, not because they had any involvement.

The corpus, the fetch tooling with its per-document hashes, the assessment runner, the deterministic claim resolver, its self-test, and the corpus validator are all committed. What stays proprietary is the calibrated assessment methodology: the ratified grading rules, worked reference rows, thresholds, and axis weighting. The instrument used here was built only from published 4+1 material and carries none of it, which is why nothing it produced can be read as an assessment.

Download the raw lab detail (Markdown)