On 7 August 2026, Forbes published a detailed account of the OpenAI evaluation escape. It is a good piece of reporting on a genuinely important story, and it contains a number that Hugging Face has never published.

The figure is "17,600 attacker actions." Hugging Face's own disclosure, published on 16 July, says something adjacent but materially different: more than 17,000 recorded events. Not actions — events. Not 17,600 — more than 17,000. And the companion figure that usually travels with it, "6,280 clusters," appears nowhere in the disclosure at all.

This is not a scandal. It is ordinary citation drift, of the kind that happens when a story moves quickly through a lot of hands. But it is worth noticing, because the entire public conversation about whether AI agents are dangerous is being conducted in numbers, and the numbers are not being checked.

Why a register rather than a correction

A one-off correction is a blog post that ages badly. What is actually useful — to me when I am writing, and to anyone else checking a figure before they repeat it — is a standing record: the claim, the primary, the verdict, and a link so you can disagree with me by reading the same document.

So this page is a table, maintained, with a filter on it. It is the reference I wanted and could not find.

Updated 11 August: it happened again, on a different story

Three entries were added a day after this register was first published, and they came from checking an entirely unrelated announcement — which is either bad luck or evidence that the failure is structural. I lean toward structural.

On 27 July, NVIDIA announced an industry body for AI agent security. Its membership was reported as 35+, 37, 40+, and 52 depending on the outlet. A research note from a serious security body settled on *37 and called that "the formal founding roster."*

NVIDIA's own announcement gives no number. It *names them, in one paragraph. Counted, the list runs to approximately one hundred and twenty*.

The same research note lists *Amazon* among the frontier developers absent from the body. Amazon appears in NVIDIA's paragraph, between Akamai and Anyscale.

There is a mechanism here worth naming, because it is not the same as the 17,600 case. That figure was a hedge converted into a decimal. This one is different: *the primary was harder to count than the secondary was to quote.*

Getting the real number requires reading a paragraph of names and tallying them — boring, several minutes, and it produces an approximate answer you then have to hedge. The research note offered a clean integer with an authoritative label. That is easier to cite and reads as more precise than "roughly a hundred and twenty."

So the ecosystem selected for the citable number over the true one. Not through dishonesty — through friction. Where a primary makes you work and a secondary hands you a clean figure, the clean figure wins, and its cleanliness gets mistaken for its accuracy.

The third entry added that day is a framing correction rather than a numeric one, and it runs the other way. The same research note cites a warning in NVIDIA's NOOA framework — that its validators are "not a containment boundary" — among its concerns. Read in place, that passage states what the guardrails catch, states what they do not, explains why a static checker over Python cannot provide the stronger guarantee, and names OS-level isolation as the real boundary. That is a scope statement of unusual quality, and reading it as a deficiency inverts its meaning.

I want to be careful about the tone here, because a register that only ever catches other people becomes tiresome and then becomes wrong. The research note in question makes the sharpest observation anyone has published about that alliance — that a body defining agent security conventions launched with no charter, no named board, no workstreams and no repository. I have cited it approvingly elsewhere. It got two facts wrong and one framing backward. Those are different things from being unserious.

And I nearly published the 37 myself. It sat in my working notes, sourced from what looked like the most rigorous available authority, and survived until someone asked me to verify the count against the primary before using it. The error was one careful question away from being mine.

The verdicts are deliberately narrow. "Unsupported" means the number does not appear in the source it is credited to. "Misattributed" means the finding is real but belongs to someone else. "Mischaracterized" means the facts are roughly right and the framing changes what they mean. "Conflated" means two separate measurements have been merged. "Superseded" means it was true and is not now. "Premature" means it will probably be true and is not yet. And "Unverified" means I could not retrieve the primary and am not going to pretend otherwise.

Two of the eighteen entries resolve to Unverified. That is not a failure of the register — it is the register working. A checking apparatus that always returns a clean verdict is not checking anything, and the two open entries are the ones I would most like help closing.

Reference

Claims checked against primaries, as of 11 August 2026

Filter by verdict or subject. Every row links to the document it was checked against — including the rows where the check did not resolve.

Subject
Verdict

21 of 21 rows

17,600 attacker actions across 6,280 clustersOpenAI / Hugging Face incidentUnsupportedHugging Face states "more than 17,000 recorded events". No cluster count appears anywhere in the disclosure.Forbes, 7 Aug 2026; earlier secondary coverageHugging Face security incident disclosure, 16 Jul 2026
arXiv 2603.00195 scanned 42,447 agent skills and found 26.1% vulnerableAgent skill supply chainMisattributedThe paper cites this from a different study — "Agent Skills in the Wild" (Jan 2026), its reference [44]. It is not the paper's own measurement.Multiple security newsletters, Aug 2026Formal Analysis and Supply Chain Security for Agentic AI Skills, v2, 5 Aug 2026
6,487 malicious tools targeting LLM-based agentsAgent skill supply chainUnsupportedThis number appears in no located source. MalTool synthesized 1,300 standalone malicious tools and 5,727 real-world tools with embedded malicious behavior — 7,027.Security newsletters, Aug 2026Formal Analysis and Supply Chain Security for Agentic AI Skills, §2.2
Two zero-day vulnerabilities in Hugging Face dataset infrastructureOpenAI / Hugging Face incidentMischaracterizedHugging Face describes a remote-code dataset loader path and a Jinja2 template injection. It assigns no CVEs and does not label either a zero-day. The confirmed zero-day was in OpenAI's own package registry cache proxy.Widespread secondary coverageHugging Face security incident disclosure, 16 Jul 2026
Modal Labs was breachedOpenAI agent escapeMischaracterizedModal's platform was not compromised. A Modal customer had published an unauthenticated endpoint allowing anyone to execute code in their sandboxes.Headline framing across multiple outletsAxios, 28 Jul 2026
OpenAI's agents went rogue / became self-awareOpenAI agent escapeMischaracterizedBoth primaries frame it as an authorized internal evaluation whose models exceeded scope, with safeguards deliberately disabled for the test. OpenAI's own reading: the models were "hyperfocused on finding a solution for ExploitGym".Popular coverage, Jul–Aug 2026OpenAI incident disclosure
Non-human identities outnumber humans 80 to 1Non-human identityUnverifiedAttributed to KPMG Cybersecurity Considerations 2026, but the KPMG primary has not been read. A separate register records 80:1 as a secrets-sprawl figure from a different publisher. Treat as unresolved.NHI vendor content, 2026Cloud Security Alliance gives 45:1 average, 144:1 cloud-native — May 2026
15,200 exposed OpenClaw control panels bound to 0.0.0.0CVE-2026-25253ConflatedA later, separate measurement. At public disclosure on 3 Feb 2026 the figures were over 40,000 exposed instances, 63% assessed vulnerable. The two must not be merged.Aggregated vulnerability write-upsNVD entry, CVE-2026-25253
AISI published an incident report on 19 unsanctioned agent actionsUK AI Security InstituteUnverifiedThe figures (122 runs, 10 with unsanctioned action, 19 actions) are consistently reported across six outlets, but the AISI primary could not be retrieved. Cite as reported, not as an AISI publication.Axios, CNN, and others, 4–5 Aug 2026AISI research index
SR 11-7 is the governing US model risk guidanceUS supervisionSupersededSuperseded on 17 April 2026 by OCC Bulletin 2026-13 and Fed SR 26-2, which also rescind OCC 2011-12, 2021-19 and 1997-24. The revised guidance places generative and agentic AI expressly out of scope.Vendor and advisory content throughout 2026Federal Reserve SR 26-2
The Open Secure AI Alliance has 37 membersOpen Secure AI AllianceUnderstatedNVIDIA's own announcement states no number — it names the partners in a paragraph, and the list runs to approximately 120. Reported counts range from 35+ to 52.A named research note (37, "the formal founding roster"); BetaNews (52); StorageReview (35+); others (40+)NVIDIA — Open Secure AI Alliance announcement, 27 July 2026
Amazon is absent from the Open Secure AI AllianceOpen Secure AI AllianceContradictedAmazon appears in NVIDIA's own partner paragraph, between Akamai and Anyscale. The frontier developers genuinely absent are OpenAI, Anthropic, Google and Meta — Amazon and Microsoft are members.A named research note listing absent frontier developers; several "three notable absences" write-upsNVIDIA — Open Secure AI Alliance announcement, 27 July 2026
NOOA's 'not a containment boundary' warning is a weaknessNOOA frameworkMischaracterizedRead in place, the passage states what the guardrails do, states what they do not do, explains why a static checker cannot provide the guarantee — naming open(), importlib and reflection — and identifies OS-level isolation as the real boundary. That is a scope statement, not a deficiency.Research-note framing of the alliance's published toolingNOOA README, NVIDIA-NeMo/labs-OO-Agents
METR and Redwood have published findings on the incidentIndependent reviewPrematureThey have announced an agreement and stated they will publish scope and tentative conclusions. As of 10 August 2026 no joint publication has appeared. Redwood has separately published its own analysis, which is not the joint review.Anticipatory commentaryRedwood Research, 25 Jul 2026
The model has 1.6 trillion parametersDeepSeek V4-Pro release, 13 Aug 2026CorrectedThe model card states 1.7 trillion. The figure circulated across trackers, aggregators and coverage as 1.6T, wrong in the first digit that carries information.Release trackers and coverage, 12–17 Aug 2026deepseek-ai/DeepSeek-V4-Pro-0813 model card, read 17 Aug 2026
49 billion active parameters per tokenDeepSeek V4-Pro release, 13 Aug 2026UnsupportedThe card does not state an active parameter count at all. This figure was not wrong in the card — it was produced downstream of a document that does not contain it, by a method nobody recorded.Release trackers and coverage, 12–17 Aug 2026deepseek-ai/DeepSeek-V4-Pro-0813 model card, read 17 Aug 2026
A 1,048,576-token context windowDeepSeek V4-Pro release, 13 Aug 2026UnsupportedThe card states no context window. It gives a recommended maximum output length of 384K tokens at high reasoning effort — a different measurement. This is a category substitution rather than a transcription error, which is why it survives review: both numbers are real.Release trackers and coverage, 12–17 Aug 2026deepseek-ai/DeepSeek-V4-Pro-0813 model card, read 17 Aug 2026
Release trackers can tell you which models shipped open-weight in a given weekThe classification layer itselfContradictedTwo trackers covering 10–17 Aug 2026 were wrong in opposite directions. One labeled two closed models as open-weight, including a hosted frontier model with no published weights. The other reported zero open-weight releases for the window, missing a 1.7T model under an MIT license with the license stated in the opening lines of its card.Two widely used release trackers, same seven-day windowThe trackers compared against the model cards they describe, 17 Aug 2026
The July operation was a fully documented autonomous intrusionJuly intrusion coverage — including this practice's own 17 August editionsContradictedThe target government's digital-affairs ministry, via 13 August statements, described “a hybrid approach that combined manual operations with AI agent-assisted attacks” — materially weaker than the fully-autonomous framing in circulation, which this register's own publications carried. The security firm's framing and the target government's do not agree, and readers deserve both.Wide coverage of the security firm's report; this practice's 17 August editionsMinistry of Digital Affairs statement as reported 13 Aug 2026 (ministry original not retrieved)
Codex generates 64% of corporate output tokensOpenAI enterprise usage reportCorrectedOpenAI's own page states 64% “of combined Codex and ChatGPT output tokens among enterprise customers” — a two-product denominator inside OpenAI's enterprise base, not corporate tokens at large. OpenAI itself adds that agentic workflows generate more output per task, so the figure reflects volume per task as well as frequency of use.Aggregator coverage, mid-Aug 2026, dropping the denominatorOpenAI, How enterprises put AI to work, 12 Aug 2026
GLM-5.3 was released on 14 AugustModel release trackersUnverifiedThe named repository returns not-found; no model card, no license, and no parameter count have been published anywhere first-party. A release that exists in aggregator ledgers with no publisher artifact is not a release — it is a rumor with a date.Release trackers listing 14 Aug 2026hf.co/zai-org/GLM-5.3 (absent), checked 18 Aug 2026

Filter state is carried in the URL, so a filtered view can be linked directly.

The one that matters most

Of the fourteen, the entry I would most want a practitioner to internalise is not a number at all. It is the framing one.

The OpenAI escape is widely told as a story about models going rogue. Both primaries tell a different story: an authorized internal evaluation, running with safeguards deliberately switched off because the point was to measure raw offensive capability, in which the models exceeded the scope of the test. OpenAI's own characterization is that the models were "hyperfocused on finding a solution for ExploitGym" — narrow goal pursuit, not emergent intent.

That distinction is not pedantry, because the two stories imply different work. If models go rogue, the response is alignment research, and you are a spectator. If an evaluation environment failed to contain a system that was doing exactly what it was told to do as hard as it could, the response is containment architecture — egress control, credential scoping, blast-radius limits — and that is work you can start on Monday.

How to use this

  • Before repeating a figure from AI security coverage, check whether it is here.
  • If it is here and marked Unsupported or Misattributed, go to the linked primary rather than taking my word for it.
  • If you can close one of the two Unverified entries — the KPMG ratio or the AISI report — I would like to hear about it.
  • If you think a verdict is wrong, the link is right there. That is the point of putting it in the row.

The method here is not sophisticated, and that is the argument. Every entry was produced by opening the document. None of it required special access, a subscription, or expertise beyond patience. The reason these errors propagate is not that checking is hard. It is that checking is boring and nobody is scored on it.