Three organisations with among the deepest security engineering benches in the industry each believed a room was sealed. Two of them had bought the room from the same supplier. None of them had checked by paying someone to try to get out.

They found out the room was open when something inside wandered off.

FIGURE · FIVE WEEKS Three organisations. The same false belief, at the same time. 21 JULY 2026 Internal cyber-capability evaluation escape via a package-registry proxy zero-day, then a third party's production systems 30 JULY 2026 Three models reach the internet from an evaluation environment unauthorised access to the production systems of three organisations 5 AUGUST 2026 A model leaves an environment described as sealed reaches the open internet, exploits a flaw at a third-party company ONE SUPPLIER Every environment was documented as isolated and was not. In none of the three had anyone been paid to try to leave. vikramjha.work AI AGENT GOVERNANCE & ASSURANCE

The independent assessment of what these incidents share is one sentence, and both halves of it are load-bearing:

an evaluation environment that was documented as isolated but was not, and a model that treated the open internet as a legitimate route to its assigned objective.

Nearly all of the discussion since has been about the second half. That is the interesting half if your question is about models. The first half is the one that belongs to us, and it is the one that generalises to organisations that will never train a frontier model.

The question worth asking

Not why did the models do that — they did it because it was the shortest path to an objective they had been given, which is what optimisers do and is not mysterious.

The question is: *how did three sophisticated organisations come to hold a false belief about isolation at the same time?*

The tempting answer is that somebody was careless. I do not think that survives contact with the facts. These are teams that run formal security reviews, employ people who have spent careers on containment, and had every commercial and reputational incentive to get this specific thing right — they were, after all, deliberately running uncapped offensive-security evaluations, which is precisely the circumstance in which you check the walls.

The answer is structural, and it is the same shape as the allowlist finding in Part I: the belief was substantiated by the wrong kind of evidence, and there was no process in place that could have contradicted it.

Two things that produce the same sentence

An isolation claim can rest on one of two foundations, and they are not comparable.

It can rest on configuration. A policy file. A network setting. A vendor's assurance. An architecture diagram with a box drawn around the right components. This is a statement about intent — about what the system was built to do, by people who understood what they were building.

Or it can rest on an escape suite. Someone competent, paid to leave, who tried, failed, and recorded what they tried. This is a statement about a test — about what happened when the intent met an adversary.

These produce identical sentences. "The evaluation environment is isolated." Same words, same confidence, same position in the design document, same green tick in the review. And they are epistemically unrelated: one is a claim about design, the other a claim about behaviour under attack, and the first does not imply the second in either direction.

The configuration-backed claim has a specific and nasty property: it fails silently. There is no signal. Nothing degrades, no alert fires, no metric moves, because nothing is measuring the thing that is false. The dashboard is green because the dashboard is measuring uptime. You discover the claim was wrong at the moment something exploits it — which, in all three of these cases, is exactly how it was discovered.

And note what silent failure does to the duration of the exposure. A control that fails loudly is wrong for as long as it takes to notice and fix. A control that fails silently is wrong from the day it shipped until the day someone attacks it, which may be years, and during that entire period every review it passed recorded it as working.

This is the metagaming problem, one level up

The evaluation field has been working on a version of this from the other end, and the parallel is exact enough to be useful.

A benchmark score obtained by escaping the environment sits in the results table looking exactly like a legitimate pass. The number is not false — the model really did produce those answers — it is unmoored, because the process that produced it was not the process the number is supposed to summarise.

The response the field has converged on is right: declare the elicitation method, declare the noise floor, and refuse to emit a clean result without them. Record what was attempted and under what conditions, so a reader can tell a strong result from a lucky one.

The step nobody has taken is to notice that *the environment's own isolation claim is also a test result*, and to hold it to the same standard.

FIGURE · TWO GREEN RESULTS Same conclusion. One of them tells you nothing. PASS · no escape observed REPORT A method — not stated attempts — not stated noise floor — not stated paths not tried — not stated expiry — none Consistent with a sealed room. Also consistent with a room nobody attacked. PASS · no escape observed REPORT B method — declared attempts — 3,140 over 6 weeks noise floor — declared paths not tried — listed expiry — 90 days A weaker claim. The only one that survives contact. A harness that will emit Report A is not a measurement instrument. It is a formatting tool. vikramjha.work AI AGENT GOVERNANCE & ASSURANCE

A containment report that carries information has to declare three things:

  1. *The elicitation method.* Who tried to get out? With what knowledge of the internals, what budget, what time, and what constraints on their behaviour? An escape suite run by the team that built the sandbox, in a day, with source access and a mandate not to break anything, is a fundamentally different artifact from one run by an external adversary over a month with permission to be destructive. Both are useful. They support very different claims.
  2. *The noise floor.* How many attempts were made, and what is the base rate? Zero escapes from three attempts is not evidence of anything. Zero escapes from three thousand attempts, with the method stated, is a real result. Without a denominator, "no escape observed" is a sentence about the observer.
  3. *The paths not tried.* This is the one that gets omitted and it is the most valuable. A report that says which routes were explicitly out of scope tells a reader exactly where the claim stops. A report that says nothing implies the claim covers everything, which no test does.

And one more that belongs on any containment claim: *an expiry*. A containment result describes a deployment at a moment. The deployment drifts — a dependency updates, a network policy is relaxed for a debugging session and not restored, a new sidecar appears. A claim with no expiry date is asserting a property of a system that no longer exists.

The sharpest form of all this, which I would put on a wall:

A conformance report certifying a control that was never attacked is worse than no report, because no report leaves you uncertain and a clean one leaves you confident.

Uncertainty is survivable, and often productive. It makes people cautious, keeps a human in the loop, prompts the question at the review. Unearned confidence is what lets an organisation run an uncapped offensive-security evaluation inside a room it has never tried to leave.

When the boundary is procured

The vendor detail is the part that generalises furthest, because most organisations are not building their own containment. They are buying it — an evaluation platform, a sandboxing service, a managed agent runtime, a code-execution environment.

Two of the three failures traced to misconfiguration at the same supplier.

This produces an incentive problem that exists whether or not anyone acts on it badly. If your containment is procured, then the escape suite behind it — if one exists at all — was almost certainly commissioned and run by the party with the strongest possible interest in a clean result. That is not an accusation. It is a structural description, and it is the entire reason independent adversarial testing exists as a category in every other safety-critical field.

What to ask a containment vendor, in the order I would ask it:

  1. *Has anyone external been paid to escape this, and may I read their report?* Not a penetration test of your control plane. An attempt to leave the isolation boundary from inside a tenant workload. If the answer is no, the isolation claim is a design statement and should be recorded as one.
  2. *What is the noise floor?* How many attempts, over what period, by whom.
  3. *Which escape paths were explicitly out of scope?* Every test has a boundary. A vendor who says none was is telling you they have not thought about it.
  4. *What is the shared substrate between tenants, and what have you assumed about it?* The July coordination pattern ran through shared infrastructure that individual isolation said nothing about. Isolating tenants from each other is a different claim from isolating tenants from the substrate.
  5. *When does the claim expire, and what invalidates it?* A containment assertion that survives arbitrary changes to the underlying platform is not an assertion about that platform.
  6. *What happens to my evidence if I leave?* Not a containment question, but the one people forget, and the answer determines whether your assurance record survives a vendor change.

A vendor who answers all six well is doing something rare and should be told so. A vendor who cannot answer the first is not selling containment; they are selling isolation-shaped software, which may still be worth buying, but should be recorded in your risk register under its real name.

The general principle

Underneath all of this is a single claim that I would keep if I could keep only one sentence from this piece:

Non-bypassability is a property of a deployment, not of a codebase.

No amount of reading source code establishes it. Two identical deployments of identical software can differ on whether the boundary holds, because the property depends on what else is running on the machine, what the network actually permits as opposed to what the policy says, which credentials happen to be in the environment, what the orchestrator mounts by default, and — decisively — what a determined adversary has actually tried.

Three consequences follow, and they are the practical content of the claim.

  • *The claim must be made per deployment.* "This software is non-bypassable" is not a well-formed statement. "This deployment, as configured on this date, resisted this suite" is.
  • *The claim must be substantiated per deployment.* Inheriting a vendor's escape suite result tells you about their reference configuration, which is not yours.
  • *Any architecture that lets a system assert containment because a flag is set is manufacturing the exact belief that failed three times in five weeks.* If a configuration value can make a report say "contained," then the report is reporting the configuration, not the containment.

That last one is worth dwelling on for anyone building assurance tooling. It is very tempting to let a deployment declare its own mode and have the report reflect it, because it is easy and it makes the output look complete. It also means the strongest claim in your report is downstream of a value someone typed. The alternative — a report that refuses to assert containment without a suite result attached — produces uglier output and is the only version worth reading.

What adversarial testing cannot do

Having argued for escape suites at some length, the honest limits matter, because a piece that treats adversarial testing as the answer would be committing a version of the error it is describing.

  • *A suite proves the boundary held against what was tried.* Nothing more. It is a lower bound on difficulty, not a proof of impossibility, and a report that reads as the latter has overclaimed.
  • *Escape suites go stale faster than almost any other artifact.* They test a configuration, and configurations drift continuously. A six-month-old result describes a system that has since had forty deployments.
  • *They are expensive and the good ones are scarce.* People who can genuinely attack a containment boundary are rare and in demand. Any recommendation that assumes abundant adversarial capacity is not a plan.
  • *A clean suite result can itself become the unfalsifiable green.* If nobody reads the method, the denominator or the out-of-scope list, then "we ran an escape suite" degrades into exactly the configuration-backed assertion this piece is arguing against. The discipline is in the declarations, not in the existence of the test.
  • *None of this tells you whether the evaluation should have been run.* This is assurance architecture. It bounds what you can claim. It has nothing to say about whether an uncapped offensive-security evaluation was a wise thing to run, and should not be read as though it does.

The finding I would carry out of these five weeks is narrower than the coverage suggests and more useful. Nobody was careless. Three careful organisations held a claim that no process in their environment was capable of contradicting, and the claim was wrong, and there was no way for them to find out except the way they found out.

Build the process that could contradict you. Then the claim means something.