Inherit the boundary. Own the gate.
By Keith Townsend
Can a team of analysts run my vendor assessment without their evidence mixing? Kamiwaza sponsored this lab to test that on its enterprise AI platform. I set up two workrooms, Kamiwaza's isolated workspaces, each holding one vendor's public documents: Dell in one, Supermicro in the other. Each workroom had its own AI agent running the same eight-layer assessment. The test fails if an agent pulls evidence from the other workroom, or scores a vendor on documents nobody approved, while the platform reports that everything worked.
The isolation held. Asked again and again about the vendor in the other workroom, retrieval returned 241 results and none came from next door. What Kamiwaza can't do is stop the person running the assessment from also changing the documents it reads, because no role separates the two. So the check that the evidence is the approved set is code you write yourself. That is the title in plain terms: take the isolation the platform provides, and keep the evidence check. Measured on one deployment, one release, two workrooms, one assessment workload and one pinned model, not projected to tenant scale.
Do Inherit the evidence boundary instead of writing one. Isolation is a property of the workroom, not of retrieval code your team maintains: 51 adversarial cross-vendor searches returned 241 hits and not one came from the other workroom, with the agent asking for the absent vendor by name throughout.
Do Own the gate. Build a deterministic corpus validator and run it before scoring rather than as a report afterward. It reads the corpus itself, so the party who contaminated it cannot delete their way out of the finding. This is your work on any platform.
Don’t Do not expect a role to express "can run the instrument, cannot touch the evidence." Both arrive on the same role and the vocabulary is a closed set. Price the process, not a permission.
Don’t Do not trust a status field on this deployment. A 201 has meant no index, a 200 no create, DEPLOYED a dead upstream, and terminal_outcome "indexed" zero vectors. Verify by reading state back.
Kamiwaza sponsored this lab. They supplied a demo deployment and an administrative token, and they did not see the production instrument or choose the vendors. Kamiwaza reviewed the findings before publication for factual errors. Dell and Supermicro appear only as corpus material from their published press releases and neither is scored; Dell is part of the practice's vendor network. Kamiwaza was not scored either. The findings below include the ones they would not have chosen.
Every vendor row I publish in the 4+1 AI Infrastructure model is my judgment applied to evidence I picked. That doesn't scale past me. The obvious fix is a team of analysts, one vendor each, all running the same assessment. So why haven't I built it? Because if each analyst controls which documents their assessment reads, two rows scored against two different sets of documents can't be compared, and comparison is the reason a vendor matrix exists.
So the question for this lab was narrow. Can a platform keep each analyst's AI agent inside the documents I approved for that analyst, without me writing the isolation code myself?
Kamiwaza sponsored the test, and I ran it on their platform. The unit of isolation there is a workroom: a workspace that holds its own documents, its own agent and its own members. I built two. One held eight public press releases from Dell. The other held eight from Supermicro. Each workroom got its own Kaizen agent, Kamiwaza's built-in agent, running the same eight-layer assessment I wrote from published 4+1 material. I also named a third vendor in the questions and loaded its documents nowhere, to see whether the model would make things up. Nine assessment runs in total. A script, not my reading, checked every claim: does the cited document exist in that workroom, and does the quoted text appear in it word for word?
Did the isolation hold?
Yes. I asked each workroom to assess the vendor it had no documents for, by name, over and over. The answers alone can't prove isolation, because an agent can refuse to answer and still have pulled the other vendor's documents. So I counted what the search layer returned. Across 51 searches about the other vendor, 241 results came back, and every one came from the asking workroom's own documents.
The claims held up too. All 42 of 42 claims in the same-vendor runs cited a document in their own workroom with a verbatim quote. When evidence was missing, the agent said so: 32 of 32 cells in the cross-vendor runs came back as insufficient evidence, and across 75 cells, including the vendor loaded nowhere, the model never invented a capability it couldn't cite.
The first number matters more than the others. The isolation is enforced where documents are fetched, before the model sees anything, so it doesn't depend on which model is serving or how well it behaves. That is also the part you can't confirm by reading documentation. You have to query it and count.
What do you give up for it?
I score this with the Decision Authority Placement Model (DAPM), which asks who holds a decision and whether you could take it elsewhere. Retained means you keep it. Ceded means you depend on the vendor's own way of doing it, and moving would mean rebuilding. In the 4+1 model, Layer 1B is retrieval, the part that finds and returns documents, and Layer 2C is where the agent and its governance live. Both are Ceded to Kamiwaza. My published assessment rated them that way in July from documentation, and this lab doesn't change the rating. What it adds is evidence that the thing you depend on actually holds under pressure.
Layer 1C, deciding which documents count as evidence, stays Retained. That's your policy, not the vendor's. So the next question was whether the platform lets you express it.
Is there a role for "run it, don't edit it"?
No. The roles are owner, editor and viewer, and that's the whole list. The viewer role is locked down properly. I tried four ways to write to the documents as a viewer, including a direct write to the vector store that skips the upload pipeline, and all four were refused. But a viewer also can't start a conversation with the agent, so a viewer can't run the assessment. Whoever runs the assessment has to be an editor, and an editor can change the documents it reads.
That isn't a Kamiwaza defect. It's ordinary access control, and the platform is clear about what its roles do. It also can't be solved with a permission, because the capability you'd want to withhold is the one you have to grant.
What do you have to build?
A check that lists every document in the workroom and compares it to the set you approved. In the lab it caught a planted document. Where you run it decides whether it protects anything. Run it after scoring and it tells you a score was already produced from bad evidence. Run it before scoring and bad evidence never gets scored. Because the check reads the documents themselves, the person who added the bad one can't hide it without also removing it.
Recovery is a rebuild. There's no way to delete a single document, so you delete the collection and reload the approved set. It took about four minutes per workroom for eight documents.
One practical warning for anyone building on this. Don't trust status responses on a platform you didn't build. An upload returned 201 and indexed nothing. A load reported success into a collection the agent never reads. Both times the fix was to count what was actually stored.
What this didn't test
It wasn't a security test. Every account was authorized and acting in good faith, and nobody tried to break in. It was one deployment, one release, two workrooms and one kind of work, with one run per direction rather than a repeated trial. I didn't test whether an update to the assessment reaches every workroom, and the model's refusal behavior belongs to that model, not to Kamiwaza.
So the verdict is the title. Take the isolation where the platform enforces it at retrieval, and stop rebuilding it in your own code. Then write the evidence check yourself and run it before anything gets scored, because no platform role will do that job for you.
The numbers
Where each layer belongs
Read for portability. Could you take this layer elsewhere without rebuilding? How to read this table
| Layer | Placement |
|---|---|
Layer 1B · Retrieval Context Management & Retrieval The retrieval path, the workroom binding and the isolation model are Kamiwaza's own, and an enterprise can't lift them out and run them on another substrate without rebuilding. The canon places every 1B component at Ceded, and nothing in this lab moves it. What the lab adds is the return on that cession, measured rather than assumed: asked 51 times about the vendor in the neighbouring workroom, retrieval returned 241 hits and every one came from the asking workroom's own corpus. The boundary is enforced at retrieval, so it doesn't depend on which model is serving. | Ceded |
Layer 1C · Pipelines Data Movement & Pipelines The canon marks 1C a gap and Enterprise Responsibility, so Kamiwaza provides nothing here and the layer is Retained. The curation policy, the vetted manifest and the corpus validator that checks against it are the enterprise's own code and would run on any platform. What the enterprise can't do on this platform is express the separation it most wants, because no role distinguishes running the instrument from writing the corpus. Read-scoped and write-scoped tokens produced identical outcomes, so token scope isn't an authorization boundary here. | Retained |
Layer 2C · Reasoning Agentic Infrastructure — The Reasoning Plane ReBAC enforcement and agent lifecycle governance are proprietary to Kamiwaza, and the canon already places them at Ceded. The agent is an artifact whose instructions and model binding are fixed at creation and bound to its workroom: a principal holding editor in one workroom only, pointed directly at the other workroom's application URL, sees zero agents. The role vocabulary is Kamiwaza's too. WorkroomRole is a closed enum, so the intent an enterprise most wants to state at this layer is one it can't. | Ceded |
Layer 3 (+1) · Applications AI Application Layer — The Value Plane The Kaizen agent and App Garden are Kamiwaza's opinionated application, and its access semantics can't be lifted out, so the layer is Ceded. The canon places those components at Ceded and the deployment patterns at Delegated, and this lab found nothing that moves it. The twenty-three tools the shipped agent carries come from the runtime image, agent-level filters had no effect, and the workroom deploy path exposes no options. That surface granted nothing a role didn't already hold, so it's a reliability cost rather than a governance one. | Ceded |
Method and disclosure
Kamiwaza sponsored this lab. They supplied a demo deployment and an administrative token, and they did not see the production instrument or choose the vendors. They received a pre-publication report carrying every finding, every engineering item and the capability request, and reviewed it for factual errors before publication.
The essay on this page was drafted by Claude Opus 5.5, an Anthropic model, from this record under the author’s voice specification, passed the deterministic post gate, and was validated by the author before publication.
Dell and Supermicro appear only as corpus material, taken from their own published press releases. Neither is assessed here and no grade in this lab says anything about either company. Dell is part of the practice's vendor network, which is disclosed here because the documents are theirs, not because they had any involvement.
The corpus, the fetch tooling with its per-document hashes, the assessment runner, the deterministic claim resolver, its self-test, and the corpus validator are all committed. What stays proprietary is the calibrated assessment methodology: the ratified grading rules, worked reference rows, thresholds, and axis weighting. The instrument used here was built only from published 4+1 material and carries none of it, which is why nothing it produced can be read as an assessment.
Placement table re-scored, September 28, 2026. Every row is now read for portability, and the table says so at the top. Earlier tables across the labs mixed that reading with the other one, and some rows rested on location, ownership or cost. No placement changed here; the notes now give each row’s reason under that reading.
The essay on this page was rendered from the lab’s frozen structured record by a model writing under this practice’s voice specification, with a deterministic gate checking every figure against the record before publication. The record is the canonical surface: where prose and record disagree, the record rules.
Download the raw lab detail (Markdown)