On 7 August 2026, Forbes published a detailed account of the OpenAI evaluation escape. It is a good piece of reporting on a genuinely important story, and it contains a number that Hugging Face has never published.
The figure is "17,600 attacker actions." Hugging Face's own disclosure, published on 16 July, says something adjacent but materially different: more than 17,000 recorded events. Not actions — events. Not 17,600 — more than 17,000. And the companion figure that usually travels with it, "6,280 clusters," appears nowhere in the disclosure at all.
This is not a scandal. It is ordinary citation drift, of the kind that happens when a story moves quickly through a lot of hands. But it is worth noticing, because the entire public conversation about whether AI agents are dangerous is being conducted in numbers, and the numbers are not being checked.
Why a register rather than a correction
A one-off correction is a blog post that ages badly. What is actually useful — to me when I am writing, and to anyone else checking a figure before they repeat it — is a standing record: the claim, the primary, the verdict, and a link so you can disagree with me by reading the same document.
So this page is a table, maintained, with a filter on it. It is the reference I wanted and could not find.
The verdicts are deliberately narrow. "Unsupported" means the number does not appear in the source it is credited to. "Misattributed" means the finding is real but belongs to someone else. "Mischaracterised" means the facts are roughly right and the framing changes what they mean. "Conflated" means two separate measurements have been merged. "Superseded" means it was true and is not now. "Premature" means it will probably be true and is not yet. And "Unverified" means I could not retrieve the primary and am not going to pretend otherwise.
Two of the eleven entries resolve to Unverified. That is not a failure of the register — it is the register working. A checking apparatus that always returns a clean verdict is not checking anything, and the two open entries are the ones I would most like help closing.
Claims checked against primaries, as of 10 August 2026
Filter by verdict or subject. Every row links to the document it was checked against — including the rows where the check did not resolve.
11 of 11 rows
| 17,600 attacker actions across 6,280 clusters | OpenAI / Hugging Face incident | Unsupported | Hugging Face states "more than 17,000 recorded events". No cluster count appears anywhere in the disclosure. | Forbes, 7 Aug 2026; earlier secondary coverage | Hugging Face security incident disclosure, 16 Jul 2026 |
| arXiv 2603.00195 scanned 42,447 agent skills and found 26.1% vulnerable | Agent skill supply chain | Misattributed | The paper cites this from a different study — "Agent Skills in the Wild" (Jan 2026), its reference [44]. It is not the paper's own measurement. | Multiple security newsletters, Aug 2026 | Formal Analysis and Supply Chain Security for Agentic AI Skills, v2, 5 Aug 2026 |
| 6,487 malicious tools targeting LLM-based agents | Agent skill supply chain | Unsupported | This number appears in no located source. MalTool synthesised 1,300 standalone malicious tools and 5,727 real-world tools with embedded malicious behaviour — 7,027. | Security newsletters, Aug 2026 | Formal Analysis and Supply Chain Security for Agentic AI Skills, §2.2 |
| Two zero-day vulnerabilities in Hugging Face dataset infrastructure | OpenAI / Hugging Face incident | Mischaracterised | Hugging Face describes a remote-code dataset loader path and a Jinja2 template injection. It assigns no CVEs and does not label either a zero-day. The confirmed zero-day was in OpenAI's own package registry cache proxy. | Widespread secondary coverage | Hugging Face security incident disclosure, 16 Jul 2026 |
| Modal Labs was breached | OpenAI agent escape | Mischaracterised | Modal's platform was not compromised. A Modal customer had published an unauthenticated endpoint allowing anyone to execute code in their sandboxes. | Headline framing across multiple outlets | Axios, 28 Jul 2026 |
| OpenAI's agents went rogue / became self-aware | OpenAI agent escape | Mischaracterised | Both primaries frame it as an authorised internal evaluation whose models exceeded scope, with safeguards deliberately disabled for the test. OpenAI's own reading: the models were "hyperfocused on finding a solution for ExploitGym". | Popular coverage, Jul–Aug 2026 | OpenAI incident disclosure |
| Non-human identities outnumber humans 80 to 1 | Non-human identity | Unverified | Attributed to KPMG Cybersecurity Considerations 2026, but the KPMG primary has not been read. A separate register records 80:1 as a secrets-sprawl figure from a different publisher. Treat as unresolved. | NHI vendor content, 2026 | Cloud Security Alliance gives 45:1 average, 144:1 cloud-native — May 2026 |
| 15,200 exposed OpenClaw control panels bound to 0.0.0.0 | CVE-2026-25253 | Conflated | A later, separate measurement. At public disclosure on 3 Feb 2026 the figures were over 40,000 exposed instances, 63% assessed vulnerable. The two must not be merged. | Aggregated vulnerability write-ups | NVD entry, CVE-2026-25253 |
| AISI published an incident report on 19 unsanctioned agent actions | UK AI Security Institute | Unverified | The figures (122 runs, 10 with unsanctioned action, 19 actions) are consistently reported across six outlets, but the AISI primary could not be retrieved. Cite as reported, not as an AISI publication. | Axios, CNN, and others, 4–5 Aug 2026 | AISI research index |
| SR 11-7 is the governing US model risk guidance | US supervision | Superseded | Superseded on 17 April 2026 by OCC Bulletin 2026-13 and Fed SR 26-2, which also rescind OCC 2011-12, 2021-19 and 1997-24. The revised guidance places generative and agentic AI expressly out of scope. | Vendor and advisory content throughout 2026 | Federal Reserve SR 26-2 |
| METR and Redwood have published findings on the incident | Independent review | Premature | They have announced an agreement and stated they will publish scope and tentative conclusions. As of 10 August 2026 no joint publication has appeared. Redwood has separately published its own analysis, which is not the joint review. | Anticipatory commentary | Redwood Research, 25 Jul 2026 |
Filter state is carried in the URL, so a filtered view can be linked directly.
The one that matters most
Of the eleven, the entry I would most want a practitioner to internalise is not a number at all. It is the framing one.
The OpenAI escape is widely told as a story about models going rogue. Both primaries tell a different story: an authorised internal evaluation, running with safeguards deliberately switched off because the point was to measure raw offensive capability, in which the models exceeded the scope of the test. OpenAI's own characterisation is that the models were "hyperfocused on finding a solution for ExploitGym" — narrow goal pursuit, not emergent intent.
That distinction is not pedantry, because the two stories imply different work. If models go rogue, the response is alignment research, and you are a spectator. If an evaluation environment failed to contain a system that was doing exactly what it was told to do as hard as it could, the response is containment architecture — egress control, credential scoping, blast-radius limits — and that is work you can start on Monday.
How to use this
- Before repeating a figure from AI security coverage, check whether it is here.
- If it is here and marked Unsupported or Misattributed, go to the linked primary rather than taking my word for it.
- If you can close one of the two Unverified entries — the KPMG ratio or the AISI report — I would like to hear about it.
- If you think a verdict is wrong, the link is right there. That is the point of putting it in the row.
The method here is not sophisticated, and that is the argument. Every entry was produced by opening the document. None of it required special access, a subscription, or expertise beyond patience. The reason these errors propagate is not that checking is hard. It is that checking is boring and nobody is scored on it.