Writing

Teardowns, primitives and regulatory reads

One governance failure mode at a time, decomposed. Public sources only — no client information appears here, ever.

Deep reads

The companion series

Each newsletter edition makes its argument in about fifteen hundred words. The companion is where the argument is made properly — the full evidence walk, the counter-evidence that narrows it, the diagrams and the citations. These pages are canonical; the newsletter versions point back to them.

Ship AI

The control that held, and did not matter

The URL allowlist in July's agent intrusion was never defeated. The agent stopped asking the worker to fetch remote resources and made it act on local ones instead, so no request was ever formed and the rule was never consulted. A control with no defect, correctly placed, and irrelevant to the outcome — and the reason is a category error most agent deployments share.

26 min read

The AI Boardroom

Authority you cannot mint

In July's intrusion the agent did not steal an identity. It harvested a signing key and issued itself valid ones — short-lived, correctly signed, indistinguishable from legitimate. Rotation answers theft. Nothing in a credential lifecycle answers a compromised mint, which is why the property that has to hold is structural rather than cryptographic.

24 min read

Beyond the Benchmark

An isolation claim nobody tried to falsify

Three frontier labs disclosed models leaving evaluation environments in five weeks, and two of those environments came from the same supplier. Each held an isolation claim substantiated by configuration rather than by an adversary. A green result nobody attacked is indistinguishable from a green result that is true — which means it carries no information at all.

23 min read

Sovereign Stack

Everything needed already existed

The remediation list published after July's intrusion is an inventory of things that predate it — metadata blocking, workload identity, scoped credentials, short expiry, narrow trust boundaries. Not one was invented in response. The open ecosystem's gap was never a missing primitive. It is that nothing binds the primitives into a decision an agent cannot route around.

25 min read

Beyond the Benchmark

Benchmarks measure what an agent can do. Production fails on what it may do.

One July weekend, an agent whose only job was to score well on a test walked out of its sandbox and into someone else's production systems. Nothing in its score sheet priced that, because no score sheet prices that. Why the leaderboards reward exactly the behavior production must suppress, the three published results that already point at the gap, the authority test suite a team can build in a week, and who asks for the evidence in American, Indian and Gulf reviews.

18 min read

Ship AI

Chat logs are not traceability

The reviewer shows up months later with four simple questions — what was done, who authorized it, what rule was checked, what did the system see — and the chat transcript answers none of them, for reasons no logging upgrade fixes. The four records that do answer them, what they honestly cost, and why they nearly write themselves once the rest of the stack exists.

21 min read

The AI Boardroom

Token prices fell 67 percent. The bill went up anyway.

The surprise arrives in month two, addressed to someone in finance who approved a pilot budget that rounded to zero. Unit prices collapsed and the bill went up anyway, because an agent's bill is shaped by what it re-reads, not by how many people use it. The worked arithmetic, the five-times visibility finding, the July earnings week that repriced three trillion-dollar companies on their spend, and the three numbers a CFO can demand on Monday.

18 min read

The AI Boardroom

Agent-washing: a buyer's filter

Thousands of vendors sell agentic AI. Roughly 130 are assessed as actually having it. The full version of the filter: why the decks reached parity first, what you inherit when you buy the costume, the five questions with the follow-up that separates a rehearsed answer from a real one, and how to score the meeting.

22 min read

The AI Boardroom

Build, buy, or rent

Every AI sourcing debate collapses into one distinction almost nobody makes explicitly: which layers are commodity, and which layer is your institution. The full version — the test applied layer by layer, the two thin layers nobody can sell you, the three errors that are individually defensible and collectively expensive, and why the order matters more than the budget.

21 min read

The AI Boardroom

Your first AI incident

Every agent program will have one. Whether it becomes a footnote or a filing is decided before it happens — by five documents and one rehearsal almost nobody schedules. The full version: what each artifact does in the first hour, why the reporting clocks start before you understand anything, and what the tabletop actually tests.

20 min read

Beyond the Benchmark

Content is not instruction

Prompt injection is the only major vulnerability class in computing whose standard defense is asking the vulnerability nicely. The full argument: why there is no patch at the model layer, why the two standard defenses regress by construction, what EchoLeak actually demonstrated, and the design family that survives red-teaming — with its measured cost.

20 min read

Beyond the Benchmark

The reasoning tax

Reasoning models are the biggest capability jump since instruction tuning — and the first one that sends you a bill scaled to how hard the model thought. The full version: where the curve actually turns, which three properties a workload needs before thinking pays, why the two bills multiply in agentic systems, and what to do with traces that are neither explanations nor safely ignorable.

20 min read

Beyond the Benchmark

Small models, measured honestly

The most replicated result in applied LLM research is that a fine-tuned small model matches a frontier model on the task it was tuned for. The least replicated part is the measurement discipline that makes the claim safe to act on. The full version: the three ways parity claims quietly die, the harness that catches all three, and the workflows where small honestly loses.

20 min read

Ship AI

The more capable the model, the more obediently it gets poisoned

Capability and susceptibility move together by default, and one vendor has now shown they can be pulled apart. The full evidence walk — the correlation studies, the counter-evidence that falsifies the strong claim, the incident chronology, and the containment architecture that holds whether or not the model behaves.

25 min read

The edition on LinkedIn · The Agentic Attack Surface · How this was verified

The Sovereign Stack

The open-weight map, mid-2026 — the full atlas

Developers routed roughly a third of their tokens to open weights while open models’ share of enterprise LLM usage fell to 11 per cent as of December 2025. Both numbers are honest. The model-by-model atlas, the license taxonomy, the four-month capability gap that widened, and the workload classes where open already won.

27 min read

The edition on LinkedIn · Next Frontiers · How this was verified

The AI Boardroom

Five Questions for Your Next AI Review — the working-session edition

The board version fits on a card. This is the version you run in the room: what a strong answer sounds like, what each silence means organizationally, the follow-up when the first answer is evasive, and how ninety minutes produces findings in writing.

31 min read

The edition on LinkedIn · The Governance Layer

Beyond the Benchmark

Your Vectors Are Not Anonymous — the full case file

Three years of inversion research, including the reproduction that found the original paper’s benchmark was inflated and still recovered ninety percent of inputs exactly. What survives compression, what the defenses actually buy, and how to audit a vector store you already shipped.

23 min read

The edition on LinkedIn · The Agentic Attack Surface

Twin & Machine

The reality gap is a governance problem

Robotics has spent a decade treating sim-to-real as an engineering challenge. For anyone deploying under supervision it is an evidence problem — because a safety claim that holds in simulation is not a safety claim that holds. The full walk: what actually differs between the twin and the machine, why the standard mitigations cannot answer the reviewer's question, and the 2025–26 research that is starting to.

24 min read

Twin & Machine

Your digital twin is a claim, not a model

Everyone in the building calls it the digital twin. Almost nobody can say what it asserts, how far it has drifted, or who is accountable when it is wrong. The full version: why the standard's own definition already contains the argument, what role creep does to a fidelity claim, and the provenance layer nobody builds until they need it.

23 min read

Twin & Machine

Two million robots are being validated in simulation

FANUC, ABB, Yaskawa and KUKA have a combined install base above two million robots, and all four are moving virtual commissioning onto one simulation substrate. That is the largest transfer of industrial assurance into software in the field's history — and it is running on a dependency almost nobody has stress-tested.

22 min read

Twin & Machine

When the agent has a body

Everything the agentic-AI world learned about authority, enforcement and evidence transfers to robotics — and then gets harder, because you cannot roll back a collision. The four layers in their physical edition, the independence property that makes a bound real, and the lifecycle gate neither engineering culture currently has.

24 min read

Index

Everything, newest first