This is the working register behind “The more capable the model, the more obediently it gets poisoned.” It is published because an assurance practice that asks clients to show their evidence should show its own. It is not a summary of the essay — it is the audit trail: what was checked, what survived, what was cut, and where the argument is thinner than the prose might suggest.

How sources were graded

Every source was placed in one of two categories before drafting began. Direct sources are peer-reviewed papers, preprints read in full, vendor system cards, or first-party specifications — material where the claim and its scope conditions could be read off the document itself. Secondary sources are analyst write-ups, security-vendor blog posts and press coverage. The rule applied throughout: no number reaches a chart or a headline on a secondary source alone. Where only a secondary source existed, the claim was either verified against a primary document before publication or cut.

One category needed a rule of its own. Vendor self-reported safety evaluations — the Sonnet 5 system card numbers in particular — are direct sources for what the vendor measured, and are not independent reproductions. The essay states this explicitly at the point of use rather than in a footnote, because the whole steelman rests on those figures.

The load-bearing measurements

Four results carry the argument. BIPIA (Yi et al., KDD 2025) supplies the correlation itself — Pearson 0.6423 across 25 models against Chatbot Arena Elo, 0.6635 on text tasks, p below 0.001 — and it is the cleanest of the four because text-task compliance is least confounded by general model competence. The same paper supplies the strongest internal check on itself: on code tasks the correlation is −0.03, effectively nil. An association that appears where competence is least confounding and vanishes where it is most confounding is behaving the way a real effect behaves, not the way an artifact does.

MCPTox (Wang et al., 2025) reproduces the pattern on live infrastructure — 45 MCP servers, 353 tools, 1,312 attack cases, 20 agents — and contributes the refusal-rate finding that reframes the whole problem. Bowen et al. supply the independent surface: 24 models from 1.5B to 72B, data poisoning rather than inference-time injection, coefficients of +0.037 to +0.063 on log-parameters at p below 0.05. The joint Anthropic, UK AI Security Institute and Alan Turing Institute study supplies the constant-cost result at roughly 250 poisoned documents.

These four are doing different jobs and the essay is careful not to let them borrow each other's authority. The scaling claim rests on Bowen et al., which measured scaling directly. The 250-document result is a constant-cost finding, not a scaling finding, and is cited only for the claim that size confers no automatic protection.

The counter-evidence, and what it forced

The thesis in its strong form — that capability forces susceptibility — is false, and three sources falsify it. Bowen et al.'s own Gemma-2 exception runs the opposite direction, with larger models more robust. The Claude Sonnet 5 system card reports browser-use injection success falling from roughly 50 percent on the previous generation to under 1 percent, alongside a capability increase. OpenAI's instruction-hierarchy work and the March 2026 IH-Challenge results show trust ordering is trainable at minimal capability cost.

This is the single most consequential thing in the register. The counter-evidence arrived after the piece was outlined and required the claim to narrow, from a statement about scaling laws to a statement about defaults: susceptibility is what you get absent deliberate instruction-hierarchy work, and the agentic surface is where that work has not landed. The narrower claim is more useful — it names a surface to inspect rather than a trend to wait out — but it is narrower, and the register records that it was the evidence that narrowed it.

The July 2026 IH-Benchmark bounds the good news in turn: across 37 models and 2,336 scenarios, compliance with intended trust ordering ranges from 98.2 percent to 20.5 percent, and strong System-over-User compliance does not predict User-over-Tool robustness. Tool output is exactly where MCP lives.

The objection conceded in the text

The strongest attack on this argument is a measurement artifact, and it is conceded in the essay's second section rather than defended against later. On agentic benchmarks a weak model scores safe partly because it fails the attacker's task the way it fails the user's. AgentDojo's figures have this shape. Some portion of the measured gradient in every study cited is incompetence rather than resistance.

The concession is load-bearing in both directions and the essay says so: safety-by-incompetence is unavailable at the capability level where deployment is worthwhile, so the artifact flattens the curve among models nobody would ship and leaves it intact among the models actually under consideration.

What was cut

Three items in the research pack did not reach the published piece, and the reasons are worth recording.

The GTG-1002 disclosure — a state actor orchestrating an intrusion through a coding agent, with a reported 80 to 90 percent of the workflow executed by the model — was cut because it is a capability-as-weapon incident reached through jailbreaking, not an injection incident. Including it would have inflated the incident count by conflating two different threat models.

A claimed finding that reasoning modes reduce injection susceptibility was cut for pointing both ways in the literature: one result shows a drop on Qwen3 8B, another reports reasoning models as more susceptible to instruction-hierarchy failures. With the evidence genuinely split, the honest move was silence rather than picking the convenient side.

A precise attack-success figure for AgentDojo was held back from the headline because it was reached secondarily. The benchmark is cited for the shape of its result — that capable models absorb most successful targeted attacks — which is what the argument actually needs.

Where this argument is weakest

Three places, stated plainly. The 250-document study covered models at or below roughly 13B parameters and a low-stakes denial-of-service backdoor; extrapolating it to frontier-scale models embedding steered behaviour is not supported by the paper and is not attempted. The Sonnet 5 figures are vendor self-reports on internal harnesses and have not been independently reproduced. And the containment controls the essay closes on — plane separation, egress control, tool-description review — are argued from published research and engineering reasoning rather than from a controlled study showing they reduce incident rates in production. CaMeL's 77-percent-with-provable-security against 84-percent-undefended is the closest thing to a measured trade, and it is one benchmark.