There is a version of the same meeting I keep hearing described, and it is never about the technology. It happens at budget time. Someone from finance asks the sponsor of the AI program a question that fits in five words — what did we get? — and the room produces adoption numbers, demo footage and a roadmap, none of which is an answer. The program does not die in that meeting. It dies in the next one, when the question is asked again and the answer has not changed.

Three research houses have now measured how often that room can answer, in different ways, and arrived at the same uncomfortable place. KPMG's Global AI Pulse for the second quarter of 2026 — 2,145 senior leaders across twenty markets — found that seven percent report established ROI from AI. MIT's Project NANDA found that roughly ninety-five percent of generative-AI pilots produced no measurable effect on profit and loss, with only about five percent scaling into anything that did. Gartner's poll of 3,400 organizations projects that more than forty percent of agentic AI programs will be canceled by the end of 2027, citing escalating costs, unclear value and inadequate risk controls.

Read separately, each of those is a failure statistic and everyone has already seen it. Read together they stop being about failure and become a description of the winners. A small minority is extracting measurable value at the same time, with the same models, at the same prices, from the same handful of laboratories as everyone else. Frontier capability is now evenly distributed and cheaply available. So whatever separates that minority, it is not the AI.

I will call them the seven percent, on KPMG’s number, while noting that the honest figure is a band rather than a point — MIT's read puts it nearer five, KPMG's nearer seven, and the true figure is probably below all three. The shorthand matters less than the question underneath it: what has to be true of a program before value becomes something you can demonstrate rather than assert?

What the three studies actually measured — and where each one is weak

Before building anything on those numbers, they deserve the scrutiny that convenient statistics rarely get.

The KPMG figure is a self-report. “Established ROI” means an executive answered yes to a survey question, and executives are not neutral referees of their own programs. My expectation is that it overstates, because the incentive is to claim a measurement you are still building. What makes the survey useful is what sits beside that number: only twenty-six percent report full, real-time visibility into what their AI systems cost to operate, while planned spend runs to a weighted average of two hundred and two million dollars over the following year. Investment is not the constraint. Instrumentation is.

The MIT work has been challenged hard, and some of the criticism lands. The sample is not random, the definition of a pilot is elastic, and the ninety-five-percent headline traveled considerably further than the methodology supports. What survives the criticism is the narrower and more useful finding: the difference between initiatives that produced an effect and those that did not correlated with integration depth, not with model choice, prompt technique or spend. That is a claim about architecture, and it is testable inside your own estate.

The Gartner projection is a forecast, which is to say an opinion with a number attached. Forecasts of cancellation rates have a poor record in both directions. What makes it worth citing is not the forty percent but the three stated causes, because those are observations about programs that already exist rather than predictions about ones that do not.

And all three share a survivorship problem nobody states plainly: they measure programs that got far enough to be counted. Initiatives that died in procurement, or were quietly renamed, or never left one business unit, are not in the denominator — so the real proportion producing measurable value is probably lower than seven percent, not higher. None of which changes the shape of the finding. When flawed instruments disagree about magnitude and agree about direction, the direction is the signal.

Same models, same prices. Seven percent can prove it paid. Three studies, three methodologies, one direction — and what separates the minority is measurement, not model choice. STUDY WHAT IT MEASURED THE NUMBER WHERE IT IS WEAK KPMG Global AI Pulse Q2 2026 · 2,145 senior leaders Leaders reporting established ROI from their AI programs 7% Self-report, and the incentive is to claim measurement you are still building. MIT Project NANDA “The GenAI Divide,” 2025 GenAI pilots producing no measurable effect on P&L 95% Methodology contested. The sample is not random, and “pilot” is an elastic word. Gartner Poll of 3,400 organizations Agentic programs projected canceled by the end of 2027 40% A forecast. The three stated causes matter more than the number attached to them. All three share a survivorship problem nobody states plainly. They count only programs that got far enough to be counted; the ones that died in procurement are not in the denominator. The true share of enterprise AI effort producing measurable value is probably lower than seven percent, not higher. The countable-outcome test, worked through Name the number. Name the person who already receives it. Name what it read before you started. ARCHETYPE ONE Internal knowledge assistant FAILS THE TEST The countable outcome is “time saved.” No baseline was ever taken, so no improvement can be shown. ARCHETYPE TWO Claims or case triage PASSES CLEANLY Cycle time, cost per claim and straight-through rate are already reported — often to a regulator. ARCHETYPE THREE Coding assistant REAL VALUE, WEAK CLAIM Change failure rate and lead time are meaningful and instrumented, but they move for many reasons. Name the number · name who already receives it · name what it read before you started vikramjha.work

Five properties, none of them purchasable

Here is what the successful minority has in common. None of it is glamorous, and none of it appears on a vendor's comparison matrix.

One: they picked workflows where value is countable. Not “productivity” — a number somebody already reports to somebody else. Claims cycle time. Cost per resolved ticket. Days sales outstanding. If a workflow's value was never measured before AI, it will not become measurable afterward, and the program joins the majority that cannot prove anything either way.

Two: they decided what the system may do before deciding what it could do. The failing majority runs capability-first: build the impressive thing, then seek permission. The minority runs authority-first: define what the system may touch, who signed off, and what happens when it is wrong — then build to that envelope. It looks slower and is faster, for reasons worth taking apart separately below.

Three: they integrate into systems of record, not alongside them. This is the MIT finding restated. Tools that live in a side window get abandoned; systems wired into the workflow's actual pipes compound. Integration depth has predicted survival better than model choice ever has, and it is the one property that gets harder rather than easier the longer you defer it.

Four: they run the economics per outcome, not per license. They can state what a resolved case costs with and without the system, including the model's own consumption — which scales with every step that re-reads its context, not with the seat count. The majority can state what they spend on AI. Those are different sentences, and only one survives a budget review: leaders reporting strong cost visibility are five times more likely to report established ROI, fifteen percent against three.

Five: they can answer for the system. When something goes wrong — and it does, in every program — the survivors can say what the system knew, what authority it acted under, and what else it could have touched. That answer keeps an incident scoped to an incident. Its absence turns one bad output into a frozen program. Gartner's “inadequate risk controls” is this property missing, observed from outside. The audit is the product.

Notice what is absent: model selection, prompt technique, agent frameworks, orchestration platforms — the things the vendor ecosystem sells hardest. The minority treats all of it as commodity, because it is, and spends its scarce attention on five things nobody sells ready-made.

The countable-outcome test, worked through three archetypes

The first property sounds like a platitude until you try to apply it, so here it is against three workflows that appear in almost every enterprise.

The internal knowledge assistant fails the test almost every time. Employees ask questions, the system answers from company documents. It is the single most-deployed enterprise AI workflow, and its countable outcome is “time saved” — measured how, against what baseline, reported to whom? Nobody was measuring how long it took to find a policy document beforehand, so nobody can measure the improvement now. It can still be worth building; it cannot be the workflow you use to prove the program works.

Claims or case triage passes cleanly. Documents arrive, get classified, routed and partially adjudicated. Cycle time per claim, cost per claim and the straight-through-processing rate — the share of claims handled with no human touch — are all already reported, often to a regulator. The baseline exists, the instrument exists, and the improvement lands in a number an executive already reads monthly. This is where the seven percent live.

The coding assistant sits between the two, and is the more instructive case. Acceptance rate and lines generated are measurable and nearly meaningless. Change failure rate and lead time to production are meaningful and already instrumented in most engineering organizations — but attributing movement in them to the assistant is genuinely hard, because they move for many reasons at once. Real value, weak attribution: a program staking its ROI claim on it is making a claim it cannot defend.

The test in one sentence: name the number, name the person who already receives it, and name what it read before you started. If any of the three is missing, you have chosen a workflow that cannot produce evidence, whatever else it produces.

Why the seven percent cluster where they do

They are not evenly distributed, and the clustering is informative. The concentration sits in insurance, banking operations, telecommunications customer operations, and parts of healthcare administration. What those have in common is not technical sophistication — several are famously behind on core systems. It is that they were already measured: regulated and operations-heavy industries spent decades building the reporting apparatus the countable-outcome test requires. They know cost per claim because a supervisor made them know it.

The corollary is uncomfortable for everyone else. In an industry with weak operational measurement, an AI program has to build the measurement before it can demonstrate the value — work that is unglamorous, slow, and nearly impossible to fund inside an AI budget. It gets skipped. Then the program cannot prove anything, and the next budget cycle draws the obvious conclusion.

There is a second clustering effect and the KPMG data puts a number on it: organizations where the chief executive is accountable for AI outcomes report established ROI at fourteen percent, against four percent where nobody at that level is. That is easy to misread as a governance platitude. What accountability at that level produces is workflow selection — a program owned by the function that already carries the number picks workflows where the number exists, because that is what its own performance review understands. A program owned by a head of AI picks workflows that demonstrate AI.

The authority-first inversion is the load-bearing property

Of the five, the second changes the most downstream and is the least intuitive, so it is worth taking apart.

Capability-first is the natural order. Someone builds something impressive, momentum accumulates, and only then does the program meet the risk function and whoever owns the system of record. At that meeting the question is what this may do in production, and the honest answer is that nobody decided — it was inherited from whatever credentials the prototype was handed so it could work at all. That meeting is where most programs die, and they die slowly, which is worse. The system does not get rejected; it gets narrowed to whatever can be approved without anyone writing down an authority, and what survives has the impressive capability removed and the integration cost retained.

Authority-first inverts the sequence at almost no cost. Before building, state what the system may do in business terms — the action classes, the value ceilings, the records in scope, who granted that, and when the grant expires. That statement is not a document for a compliance folder. It is a specification: it tells the engineering team what to enforce, at which boundary, and what to record when the boundary is tested.

The reason this makes programs faster rather than slower is that it moves the hardest conversation from month eight to week two, when changing the answer is still cheap.

It also produces the artifact the fifth property requires. If you stated the authority up front, attestation — the written proof of what was done and on whose say-so — is a matter of recording which grant applied. If you did not, attestation is a reconstruction project run under time pressure by people who are also managing an incident.

Three markets, three versions of the same gap

In North America the binding constraint is the risk function, and it has just moved. In April 2026 the federal banking agencies replaced fifteen years of model risk guidance and put generative and agentic AI expressly outside its scope, with separate guidance promised. That is a deferral rather than an exemption — every obligation attached to the underlying action still binds — but it means an American program cannot wait to be told what the control is. The ones that stall arrive at second-line review expecting a rulebook to supply the authority statement they never wrote.

India's constraint arrives through data protection instead: the DPDP Act attaches purpose limitation to personal data, which makes an agent traversing systems outside its stated purpose a live exposure independent of any AI-specific rule, and routes the question to the privacy function first. Indian programs frequently hold a well-drafted policy position and a weak enforcement story — the mirror image of the American failure. In the Gulf the constraint is scale: sovereign programs concentrate consequence, and the failure mode is political before it is operational. The compensating advantage is that much of the estate is new-build, where the five properties cost a fraction of what a retrofit costs.

The honest limits of this argument

Two things I cannot claim. The first is causation. Everything above is a correlation between properties and reported outcomes, drawn from surveys with the weaknesses already named. It is entirely possible that organizations which measure well are simply better-run, and that their AI programs succeed for the same reason their supply-chain programs do. If that is the whole story the advice is still correct and the mechanism is different — which changes what you should expect from a program that adopts the five properties without the operational discipline underneath them.

The second is the numerator. There is no clean, comparable measure of how many enterprises hold a formal AI strategy against how many can prove a return, because the surveys asking those two questions are not the same surveys and do not share a population. The gap between near-universal adoption and single-digit demonstrated return is real and consistently reported; the precise ratio is not something I would put on a board slide as a measurement.

The board question

The question that follows from all of this is not what our AI strategy is. You have one. The question is narrower and much harder to answer with a slide.

For our largest AI program, can anyone in this room state the countable outcome, the authority envelope, and what we could prove if it went wrong this quarter?

If the answer is yes, the program is probably in the minority, and the follow-up is whether the same three answers exist for the second and the third. If the answer is no, you have not discovered a failure. You have discovered which of the five properties is missing, which is more actionable than a strategy review.

The work that follows is not a purchase. Naming a countable outcome takes an afternoon with whoever already receives the number. Writing an authority envelope takes a working session with the function that owns the system of record. Enforcing it is ordinary engineering in a gateway you already operate, and recording what happened at the moment it happens is a schema decision. None of that is why the seven percent are the seven percent — they are, because somebody senior asked the authority question early enough that the answer could still shape the architecture, and because they picked a boring workflow with a real number over an impressive one without.

If you want a second opinion on which of your programs would survive that board question, that is the conversation I have most weeks.