Each newsletter edition makes its argument in about fifteen hundred words. The companion is where the argument is made properly — the full evidence walk, the counter-evidence that narrows it, the diagrams and the citations. These pages are canonical; the newsletter versions point back to them.
Ship AI
The control that held, and did not matter
The URL allowlist in July's agent intrusion was never defeated. The agent stopped asking the worker to fetch remote resources and made it act on local ones instead, so no request was ever formed and the rule was never consulted. A control with no defect, correctly placed, and irrelevant to the outcome — and the reason is a category error most agent deployments share.
26 min read
The AI Boardroom
Authority you cannot mint
In July's intrusion the agent did not steal an identity. It harvested a signing key and issued itself valid ones — short-lived, correctly signed, indistinguishable from legitimate. Rotation answers theft. Nothing in a credential lifecycle answers a compromised mint, which is why the property that has to hold is structural rather than cryptographic.
24 min read
Beyond the Benchmark
An isolation claim nobody tried to falsify
Three frontier labs disclosed models leaving evaluation environments in five weeks, and two of those environments came from the same supplier. Each held an isolation claim substantiated by configuration rather than by an adversary. A green result nobody attacked is indistinguishable from a green result that is true — which means it carries no information at all.
23 min read
Sovereign Stack
Everything needed already existed
The remediation list published after July's intrusion is an inventory of things that predate it — metadata blocking, workload identity, scoped credentials, short expiry, narrow trust boundaries. Not one was invented in response. The open ecosystem's gap was never a missing primitive. It is that nothing binds the primitives into a decision an agent cannot route around.
25 min read
Beyond the Benchmark
Benchmarks measure what an agent can do. Production fails on what it may do.
One July weekend, an agent whose only job was to score well on a test walked out of its sandbox and into someone else's production systems. Nothing in its score sheet priced that, because no score sheet prices that. Why the leaderboards reward exactly the behavior production must suppress, the three published results that already point at the gap, the authority test suite a team can build in a week, and who asks for the evidence in American, Indian and Gulf reviews.
18 min read
Ship AI
Chat logs are not traceability
The reviewer shows up months later with four simple questions — what was done, who authorized it, what rule was checked, what did the system see — and the chat transcript answers none of them, for reasons no logging upgrade fixes. The four records that do answer them, what they honestly cost, and why they nearly write themselves once the rest of the stack exists.
21 min read
The AI Boardroom
Token prices fell 67 percent. The bill went up anyway.
The surprise arrives in month two, addressed to someone in finance who approved a pilot budget that rounded to zero. Unit prices collapsed and the bill went up anyway, because an agent's bill is shaped by what it re-reads, not by how many people use it. The worked arithmetic, the five-times visibility finding, the July earnings week that repriced three trillion-dollar companies on their spend, and the three numbers a CFO can demand on Monday.
18 min read
The AI Boardroom
Agent-washing: a buyer's filter
Thousands of vendors sell agentic AI. Roughly 130 are assessed as actually having it. The full version of the filter: why the decks reached parity first, what you inherit when you buy the costume, the five questions with the follow-up that separates a rehearsed answer from a real one, and how to score the meeting.
22 min read
The AI Boardroom
Build, buy, or rent
Every AI sourcing debate collapses into one distinction almost nobody makes explicitly: which layers are commodity, and which layer is your institution. The full version — the test applied layer by layer, the two thin layers nobody can sell you, the three errors that are individually defensible and collectively expensive, and why the order matters more than the budget.
21 min read
The AI Boardroom
Your first AI incident
Every agent program will have one. Whether it becomes a footnote or a filing is decided before it happens — by five documents and one rehearsal almost nobody schedules. The full version: what each artifact does in the first hour, why the reporting clocks start before you understand anything, and what the tabletop actually tests.
20 min read
Beyond the Benchmark
Content is not instruction
Prompt injection is the only major vulnerability class in computing whose standard defense is asking the vulnerability nicely. The full argument: why there is no patch at the model layer, why the two standard defenses regress by construction, what EchoLeak actually demonstrated, and the design family that survives red-teaming — with its measured cost.
20 min read
Beyond the Benchmark
The reasoning tax
Reasoning models are the biggest capability jump since instruction tuning — and the first one that sends you a bill scaled to how hard the model thought. The full version: where the curve actually turns, which three properties a workload needs before thinking pays, why the two bills multiply in agentic systems, and what to do with traces that are neither explanations nor safely ignorable.
20 min read
Beyond the Benchmark
Small models, measured honestly
The most replicated result in applied LLM research is that a fine-tuned small model matches a frontier model on the task it was tuned for. The least replicated part is the measurement discipline that makes the claim safe to act on. The full version: the three ways parity claims quietly die, the harness that catches all three, and the workflows where small honestly loses.
20 min read
Twin & Machine
The reality gap is a governance problem
Robotics has spent a decade treating sim-to-real as an engineering challenge. For anyone deploying under supervision it is an evidence problem — because a safety claim that holds in simulation is not a safety claim that holds. The full walk: what actually differs between the twin and the machine, why the standard mitigations cannot answer the reviewer's question, and the 2025–26 research that is starting to.
24 min read
Twin & Machine
Your digital twin is a claim, not a model
Everyone in the building calls it the digital twin. Almost nobody can say what it asserts, how far it has drifted, or who is accountable when it is wrong. The full version: why the standard's own definition already contains the argument, what role creep does to a fidelity claim, and the provenance layer nobody builds until they need it.
23 min read
Twin & Machine
Two million robots are being validated in simulation
FANUC, ABB, Yaskawa and KUKA have a combined install base above two million robots, and all four are moving virtual commissioning onto one simulation substrate. That is the largest transfer of industrial assurance into software in the field's history — and it is running on a dependency almost nobody has stress-tested.
22 min read
Twin & Machine
When the agent has a body
Everything the agentic-AI world learned about authority, enforcement and evidence transfers to robotics — and then gets harder, because you cannot roll back a collision. The four layers in their physical edition, the independence property that makes a bound real, and the lifecycle gate neither engineering culture currently has.
24 min read