Three research houses measured different things in different ways and arrived at the same uncomfortable place. KPMG's Global AI Pulse for the second quarter of 2026 — 2,145 senior leaders across twenty markets, all at organisations above a hundred million dollars of revenue — found that seven percent report established ROI from AI. MIT's Project NANDA, working from a review of more than three hundred publicly disclosed AI initiatives, structured interviews at 52 organisations and a survey of 153 senior leaders, found that roughly ninety-five percent of generative-AI pilots produced no measurable effect on profit and loss, with only about five percent scaling into anything that did. Gartner's poll of 3,400 organisations projects that more than forty percent of agentic AI programmes will be cancelled by the end of 2027, citing escalating costs, unclear value and inadequate risk controls.

Read separately, each of those is a failure statistic and everyone has already seen it. Read together they stop being about failure and become a description of the winners. A small minority is extracting measurable value at the same time, with the same models, at the same prices, from the same handful of laboratories as everyone else. Frontier capability is now evenly distributed and cheaply available. So whatever separates that minority, it is not the AI.

I will call them the eight percent, which is the generous end of the band — MIT's read puts it nearer five, KPMG's nearer seven, and the true figure is probably below all three for reasons I come to shortly. The shorthand matters less than the question underneath it: what has to be true of a programme before value becomes something you can demonstrate rather than assert?

What the three studies actually measured — and where each one is weak

Before building anything on those numbers, they deserve the scrutiny that convenient statistics rarely get.

The KPMG figure is a self-report. “Established ROI” means an executive answered yes to a survey question, and executives are not neutral referees of their own programmes. My expectation is that it overstates rather than understates, because the incentive on a survey is to claim a measurement you are still building. What makes the survey more useful than most is what sits beside that number: only twenty-six percent of leaders report full, real-time visibility into what their AI systems cost to operate, while planned spend runs to a weighted average of two hundred and two million dollars over the following twelve months. Investment is not the constraint. Instrumentation is.

The MIT work has been challenged hard, and some of the criticism lands. The sample is not random, the definition of a pilot is elastic, and the ninety-five-percent headline travelled considerably further than the methodology supports. What survives the criticism is the narrower and more useful finding: the difference between initiatives that produced an effect and those that did not correlated with integration depth, not with model choice, prompt technique or spend. That is a claim about architecture, and it is testable inside your own estate.

The Gartner projection is a forecast, which is to say an opinion with a number attached. Forecasts of cancellation rates have a poor record in both directions. What makes it worth citing is not the forty percent but the three stated causes, because those are observations about programmes that already exist rather than predictions about ones that do not.

And all three share a survivorship problem that nobody states plainly: they measure programmes that got far enough to be counted. The initiatives that died in procurement, or were quietly renamed, or never left a single business unit, are not in the denominator. The real proportion of enterprise AI effort producing measurable value is probably lower than eight percent, not higher.

None of that changes the shape of the finding. Three independent instruments, three different weaknesses, one direction. When flawed instruments disagree about magnitude and agree about direction, the direction is the signal.

Same models, same prices. Seven percent can prove it paid. Three studies, three methodologies, one direction — and what separates the minority is measurement, not model choice. STUDY WHAT IT MEASURED THE NUMBER WHERE IT IS WEAK KPMG Global AI Pulse Q2 2026 · 2,145 senior leaders Leaders reporting established ROI from their AI programmes 7% Self-report, and the incentive is to claim measurement you are still building. MIT Project NANDA “The GenAI Divide,” 2025 GenAI pilots producing no measurable effect on P&L 95% Methodology contested. The sample is not random, and “pilot” is an elastic word. Gartner Poll of 3,400 organisations Agentic programmes projected cancelled by the end of 2027 40% A forecast. The three stated causes matter more than the number attached to them. All three share a survivorship problem nobody states plainly. They count only programmes that got far enough to be counted; the ones that died in procurement are not in the denominator. The true share of enterprise AI effort producing measurable value is probably lower than eight percent, not higher. The countable-outcome test, worked through Name the number. Name the person who already receives it. Name what it read before you started. ARCHETYPE ONE Internal knowledge assistant FAILS THE TEST The countable outcome is “time saved.” No baseline was ever taken, so no improvement can be shown. ARCHETYPE TWO Claims or case triage PASSES CLEANLY Cycle time, cost per claim and straight-through rate are already reported — often to a regulator. ARCHETYPE THREE Coding assistant REAL VALUE, WEAK CLAIM Change failure rate and lead time are meaningful and instrumented, but they move for many reasons. Name the number · name who already receives it · name what it read before you started vikramjha.work

Five properties, none of them purchasable

Here is what the successful minority has in common. None of it is glamorous, and none of it appears on a vendor's comparison matrix.

One: they picked workflows where value is countable. Not “productivity” — a number somebody already reports to somebody else. Claims cycle time. Cost per resolved ticket. Days sales outstanding. If a workflow's value was never measured before AI, it will not become measurable afterwards, and the programme joins the majority that cannot prove anything either way.

Two: they decided what the system may do before deciding what it could do. The failing majority runs capability-first: build the impressive thing, then seek permission. The minority runs authority-first: define what the system is allowed to touch, who signed off, and what happens when it is wrong — then build to that envelope. It looks slower. It is faster, because the approval that actually gates production value belongs to the risk function, and arriving at that meeting with answers beats arriving with a demonstration.

Three: they integrate into systems of record, not alongside them. This is the MIT finding restated. Tools that live in a side window get abandoned; systems wired into the workflow's actual pipes compound. Integration depth has predicted survival better than model choice ever has, and it is the one property that gets harder rather than easier the longer you defer it.

Four: they run the economics per outcome, not per licence. They can state what a resolved case costs with and without the system, including the model's own consumption — which scales with every step that re-reads its context, not with the seat count. The majority can state what they spend on AI. Those are different sentences, and KPMG's own data suggests only one of them survives a budget review: leaders reporting strong cost visibility are five times more likely to report established ROI, fifteen percent against three.

Five: they can answer for the system. When something goes wrong — and it does, in every programme — the survivors can say what the system knew, what authority it acted under, and what else it could have touched. That answer is what keeps an incident scoped to an incident. Its absence is what turns one bad output into a frozen programme. Gartner's “inadequate risk controls” is this property missing, observed from outside.

Notice what is absent from that list: model selection, prompt technique, agent frameworks, orchestration platforms — the things the vendor ecosystem sells hardest. The minority treats all of it as commodity, because it is, and spends its scarce organisational attention on five things nobody sells ready-made.

The countable-outcome test, worked through three archetypes

The first property sounds like a platitude until you try to apply it, so here it is against three workflows that appear in almost every enterprise.

The internal knowledge assistant fails the test almost every time. Employees ask questions, the system answers from company documents. It is the single most-deployed enterprise AI workflow. What is the countable outcome? “Time saved” — measured how, against what baseline, reported to whom? Nobody was measuring how long it took to find a policy document beforehand, so nobody can measure the improvement now. This workflow can still be worth building. It cannot be the workflow you use to prove the programme works, and treating it as such is how a great deal of AI budget has quietly evaporated.

Claims or case triage passes cleanly. Documents arrive, get classified, routed and partially adjudicated. Cycle time per claim is already reported. Cost per claim is already reported. Straight-through-processing rate is already reported, often to a regulator. The baseline exists, the instrument exists, and the improvement lands in a number an executive already reads monthly. This is where the eight percent live.

The coding assistant sits between the two, and is the more instructive case. Acceptance rate and lines generated are measurable and nearly meaningless. Change failure rate, lead time to production and defect escape rate are meaningful and already instrumented in most engineering organisations — but attributing movement in them to the assistant is genuinely hard, because they move for many reasons at once. The honest position is that this workflow has real value and weak attribution, and a programme staking its ROI claim on it is making a claim it cannot defend.

The test in one sentence: name the number, name the person who already receives it, and name what it read before you started. If any of the three is missing, you have chosen a workflow that cannot produce evidence, whatever else it produces.

Why the eight percent cluster where they do

They are not evenly distributed across the economy, and the clustering is informative. The concentration sits in insurance, banking operations, telecommunications customer operations, and parts of healthcare administration. What those have in common is not technical sophistication — several are famously behind on core systems. It is that they were already measured. Regulated and operations-heavy industries have spent decades building the reporting apparatus the countable-outcome test requires. They know cost per claim because a supervisor made them know it. They know handle time because a workforce-management system has tracked it since the 1990s.

The corollary is uncomfortable for everyone else. In an industry with weak operational measurement, an AI programme has to build the measurement before it can demonstrate the value, and that work is unglamorous, slow, and nearly impossible to fund inside an AI budget. It gets skipped. Then the programme cannot prove anything, and the next budget cycle draws the obvious conclusion.

There is a second clustering effect, and the KPMG data puts a number on it: organisations where the chief executive is accountable for AI outcomes report established ROI at fourteen percent against four percent where nobody at that level is. Read that carefully, because it is easy to misread as a governance platitude. What accountability at that level actually produces is workflow selection. A programme owned by the function that already carries the number picks workflows where the number exists, because that is the only thing its own performance review understands. A programme owned by a head of AI picks workflows that demonstrate AI.

The authority-first inversion is the load-bearing property

Of the five, the second changes the most downstream and is the least intuitive, so it is worth taking apart.

Capability-first is the natural order. Someone builds something impressive, it demonstrates well, momentum accumulates, and only then does the programme meet the risk function, security, and whoever owns the system of record. At that meeting the question is what this may do in production, and the honest answer is that nobody decided — the answer was inherited from whatever credentials the prototype was handed so it could work at all.

That meeting is where most programmes die, and they die slowly, which is worse. The system does not get rejected; it gets narrowed. Scope is cut to whatever can be approved without anyone having to write down an authority. What survives is a system with the impressive capability removed and the integration cost retained.

Authority-first inverts the sequence at almost no cost. Before building, state what the system may do in business terms — the action classes, the value ceilings, the accounts or records in scope, who granted that, and when the grant expires. That statement is not a document for a compliance folder. It is a specification: it tells the engineering team what to enforce, at which boundary, and what to record when the boundary is tested.

The reason this makes programmes faster rather than slower is that it moves the hardest conversation from month eight to week two, when changing the answer is still cheap.

It also produces, as a by-product, the artefact the fifth property requires. If you stated the authority up front, attestation is a matter of recording which grant applied. If you did not, attestation is a reconstruction project, run under time pressure, by people who are also managing an incident.

Three markets, three versions of the same gap

In North America the binding constraint is the risk function, and it has just moved. In April 2026 the federal banking agencies replaced fifteen years of model risk guidance and put generative and agentic AI expressly outside its scope, with separate guidance promised. That is a deferral rather than an exemption — every obligation attached to the underlying action still binds — but it means an American programme cannot wait to be told what the control is. The programmes that stall are the ones arriving at second-line review without an authority statement and expecting a rulebook to supply one.

In India the constraint arrives through data protection rather than model risk. The DPDP Act attaches purpose limitation to personal data, which makes an agent traversing systems outside its stated purpose a live exposure independent of any AI-specific rule. Combined with the Reserve Bank's sharpening posture on model-driven decisioning, this routes the governance question to the privacy function first. The practical effect is that Indian programmes frequently hold a well-drafted policy position and a weak technical enforcement story, which is the mirror image of the American failure.

In the Gulf the constraint is scale and visibility rather than examination. Sovereign-scale programmes concentrate consequence, and the failure mode is political before it is operational. The compensating advantage is substantial: much of the estate is new-build, and in a new-build the five properties cost a fraction of what they cost as a retrofit. Of the three markets, the Gulf has the best structural opportunity to be in the eight percent by construction rather than by correction.

The honest limits of this argument

Two things I cannot claim, and it would be dishonest to imply otherwise.

The first is causation. Everything above is a correlation between properties and reported outcomes, drawn from surveys with the weaknesses already named. It is entirely possible that organisations which measure well are simply better-run organisations, and that their AI programmes succeed for the same reason their supply-chain programmes do. If that is the whole story, the advice is still correct and the mechanism is different — which is a distinction worth holding, because it changes what you would expect from a programme that adopts the five properties without the underlying operational discipline.

The second is the numerator. There is no clean, comparable measure of how many enterprises hold a formal AI strategy against how many can prove a return, because the surveys asking those two questions are not the same surveys and do not share a population. The gap between near-universal adoption and single-digit demonstrated return is real and consistently reported across instruments; the precise ratio is not something I would put on a board slide as a measurement.

The board question

The question that follows from all of this is not what our AI strategy is. You have one. The question is narrower and much harder to answer with a slide.

For our largest AI programme, can anyone in this room state the countable outcome, the authority envelope, and what we could prove if it went wrong this quarter?

If the answer is yes, the programme is probably in the minority, and the useful follow-up is whether the same three answers exist for the second and the third. If the answer is no, you have not discovered a failure. You have discovered exactly which of the five properties is missing, which is considerably more actionable than a strategy review.

The work that follows is not a purchase. Naming a countable outcome takes an afternoon with whoever already receives the number. Writing an authority envelope takes a working session with the function that owns the system of record. Enforcing it is ordinary engineering in a gateway or service layer you already operate. Recording what happened at the moment it happens is a schema decision.

None of that is why the eight percent are the eight percent. They are the eight percent because somebody senior asked the authority question early enough that the answer could still shape the architecture — and because they were willing to pick a boring workflow with a real number attached over an impressive one without.

If you want a second opinion on which of your programmes would survive that board question, that is the conversation I have most weeks.