Gartner gave the pattern a name that stuck: agent-washing — existing chatbots, RPA suites and assistants re-badged as agentic AI without the substance. Their estimate is the part that should change how you run a vendor meeting: of the thousands of vendors claiming agentic capability, on the order of 130 offer the real thing.

Sit with the arithmetic rather than the adjective. If that ratio is even approximately right, the default agentic-AI meeting on your calendar is, statistically, a re-badging exercise. Not because the vendor is dishonest — most are not — but because the category label became commercially necessary faster than the capability became technically available, and everyone in the market responded rationally to that.

And the deck will be indistinguishable from the genuine article, because decks are the one place parity was achieved immediately.

Why the presentation reached parity first

This is worth a moment, because understanding the mechanism tells you where to look and, more usefully, where not to bother looking.

Every observable surface of an agentic system is cheap to replicate. A conversational interface is cheap. A demonstration where the system appears to reason through a multi-step task is cheap, particularly when the demo path is known in advance. Language describing goals, planning and autonomy is free. An architecture diagram with a box labelled “planner” costs one afternoon.

What is expensive is the part with no visual signature: an authority model that bounds what the system may do; an enforcement point that sits between the system's intent and your systems of record; an evidence trail that reconstructs a specific action months later; and the engineering to hold a goal across a failed step rather than throwing an exception. None of those photograph well. All of them are what you are actually buying.

Which produces the specific difficulty this filter is designed for. The gap between a washed agent and a real one is invisible in exactly the medium where buying decisions are made, and highly visible in production, eight months later, when your risk function asks a question the vendor never had to answer.

What you actually inherit

The obvious cost of buying a washed agent is wasted licence spend. That is real and it is the least of it.

The structural cost is that the risk posture attaches to what the system is sold as, not to what it does. Call a system an agent internally — in the board paper, in the programme name, in the architecture review — and you have summoned the full governance surface: the authority questions, the audit expectations, the board scrutiny, the incident-response expectations, the model-risk attention. That surface arrives whether or not the underlying system can do anything to justify it.

What you inherit when you buy the costume The deck reached parity immediately. The risk posture did not. The market thousands of vendors claiming “agentic AI” re-badged assistants · RPA suites · chatbots with a new label ~130 assessed as genuinely agentic The red sliver is to scale. So is the meeting on your calendar. What a buyer takes on Governance surface genuine agent washed agent authority questions · audit expectations · board scrutiny identical for both Delivered capability genuine agent washed agent The gap is the agent risk premium, paid for chatbot capability. The damage is not the wasted licence spend. It is that the risk posture attaches to what the system is sold as — so you get the full governance surface and a fraction of the value. And the direction of travel, from the same research house Over 40% of agentic AI projects forecast cancelled by end-2027 — escalating costs, unclear business value, inadequate risk controls. 40% of enterprise applications expected to carry task-specific agents by 2026, from under 5%.

So the buyer of a washed agent takes on all of the governance surface and a fraction of the value that would justify it. The gap between those two bars is a real premium, paid in scrutiny and in programme overhead rather than in licence fees, and it is invisible on the invoice.

There is a second-order effect that does more damage in practice, and I have watched it play out more than once. A programme that has paid the agent risk premium without receiving agent capability becomes an argument against agents inside the organisation. The scrutiny was real, the cost was real, the value was thin — and the conclusion drawn is that agentic AI does not work here, when what did not work was a re-badged workflow tool carrying an unearned label. The washing does not just waste a budget cycle. It poisons the well for the genuine build that follows.

This connects to the forecast that gets quoted at every board table: over 40 percent of agentic AI projects cancelled by end-2027, on escalating costs, unclear business value and inadequate risk controls. Read the named causes against the paragraph above. Unclear business value is exactly what a washed agent produces, and inadequate risk controls is exactly what a vendor with no authority model hands you. A meaningful share of that forecast cancellation is not a technology failure at all. It is a procurement failure with a lagging indicator.

The actual boundary: selection, not sophistication

Before the questions, the distinction they are all testing — because buyers routinely apply the wrong test and get a confident wrong answer.

The wrong test is sophistication. A system can be extraordinarily elaborate — hundreds of branches, rich natural-language handling, a beautiful interface — and be pure automation. Complexity is not the boundary, and a demo optimised to look impressive is optimised on precisely the axis that does not distinguish the two.

The boundary is whether anything is selected at run time. An agent chooses among courses of action under uncertainty. Automation executes branches someone drew in advance. A thousand pre-drawn branches is still a thousand pre-drawn branches; a system that genuinely picks an approach nobody enumerated, and can be asked what constrained that pick, is doing something categorically different.

The line is selection, not sophistication An elaborate decision tree is still automation. A simple system that chooses is not. Automation input if / then branch authored by a person path A (pre-drawn) path B (pre-drawn) path C (pre-drawn) on failure: exception “it alerts an operator” Every path enumerated in advance by a person. Nothing is selected at run time. Entirely legitimate. Frequently the right purchase. Priced differently. Sophistication does not move it across the line — a thousand branches is still a thousand branches. Agentic input + goal selects among courses of action under uncertainty what bounds this choice? ← the whole diligence question acts retry differently re-plan · escalate the goal is held across the failure The test is whether anything is chosen at run time, and whether the goal survives a failed step. Automation is fine. Automation at agent prices, carrying agent risk, is not.

The second half of the boundary is goal persistence. When a step fails, automation terminates — an exception, an alert, a handoff to a human. An agentic system holds the goal across the failure: it retries differently, re-plans, or escalates with the objective intact. This is the half buyers forget to test, and it is the half that determines whether the system does useful work in an environment that does not behave.

Both halves matter, and they are why none of this requires a benchmark. You are not measuring how good the system is. You are establishing which category it belongs to — and category questions are answerable in conversation.

Question one: show me a decision it made that wasn't scripted

Ask for a concrete production example where the system selected an approach, and ask what bounded that selection.

The answer that clears it is specific and slightly awkward to give, because it involves describing a case where the system did something the vendor did not fully specify in advance. Real answers have texture: it had three ways to resolve this, it chose the second because of a condition in the data, and here is what stopped it choosing something we would not have wanted.

The answer that reveals a re-badge is that every decision, traced back, resolves to an if-then a human authored. That is automation. Automation is fine — frequently it is the correct purchase, more reliable and cheaper to assure than an agent — but it is priced differently and it carries a different risk posture, and both of those matter.

The follow-up that separates rehearsed from real: and what would it have done if that condition had been absent? A vendor describing genuine selection can answer, because they have watched the system behave across cases. A vendor describing a scripted path either cannot answer or describes another pre-drawn branch — which is itself the answer.

Question two: what happens when the plan fails mid-way?

This tests goal persistence, and it is the question vendors are least prepared for, because failure paths are the part demos omit.

The strong answer describes the system holding its objective across the failure: retrying with a different approach, re-planning around the obstacle, or escalating with context about what it was trying to achieve and how far it got. The weak answer is that it throws an exception, or the friendlier version, it alerts an operator.

That friendlier version deserves attention because it sounds like a feature. It is an honest automation answer wearing an agent's price tag. Alerting an operator is a perfectly good behaviour; it is also what every workflow tool built in the past thirty years does when a step fails.

The follow-up: walk me through a specific failure from the past month. Not a category of failure — one incident, what the system did, and what a human had to do afterwards. This is the single most informative question in the set, because it cannot be answered from a positioning document. A vendor operating at production volume has these stories immediately. A vendor who does not will offer a hypothetical, and the substitution is easy to hear.

Question three: how do I bound its authority — and where is that enforced?

This is the question that doubles as governance diligence, and if you only get to ask one, ask this one.

Five questions, and what each answer tells you Answerable in the first meeting. None requires a proof-of-concept. The question Genuine Re-badged 1 · Show me a decision it made that wasn’t scripted Choosing among courses of action under uncertainty. A production case where it selected an approach — and what bounded the selection. Every “decision” traces to an if-then a human authored. You are buying automation. 2 · What happens when the plan fails mid-way? Holds the goal across steps: retries differently, re-plans, escalates. It throws an exception. “It alerts an operator” is an automation answer at agent price. 3 · How do I bound its authority — and where is that enforced? This one doubles as your Grants, scopes, ceilings — and an enforcement point between intent and the system of record. Role-based access control on the user — who may run the tool, not what the tool may do. No authority model = the problem, 4 · What record does it leave — would my auditor accept it? A trace of one production action: what it knew, what it evaluated, under whose grant. “Full conversation history.” Chat logs are not attestation — that is an archaeology project. 5 · What does it cost per outcome at my volume? Models cost per resolved case at production volume. Real agents consume like infrastructure — per run. Per seat, like software. Either they have not run at production volume, or they would rather you did not multiply. Question three doubles as governance diligence. A vendor with no authority model is selling you the unbounded-agent problem as a feature.

The genuine answer names grants, scopes and ceilings — what the agent may do, expressed in business terms rather than in system permissions — and, critically, an enforcement point that sits between the agent's intent and your system of record. Something that evaluates the proposed action and can refuse it.

The answer that reveals the gap is role-based access control on the user. This sounds responsible and is entirely beside the point: RBAC on the user governs who may run the tool, not what the tool may do once running. Those are different controls answering different questions, and conflating them is the most common genuine confusion in this space — not a deception, usually, but a category error the vendor has not been forced to confront.

A vendor with no authority model is handing you the unbounded-agent problem as a feature. Whatever the system's capability turns out to be, you will be building the bounding layer yourself, on their timeline, against their API. That is a very different programme from the one in the proposal, and it should be priced as one.

The follow-up: show me where a request gets refused. Not described — shown. A vendor with a real enforcement point can demonstrate a denial in about a minute, because denial is a code path they had to build and test. A vendor without one will explain the policy framework.

Question four: what record does it leave — and would my auditor accept it?

Ask to see the trace of one production action: what the system knew, what it evaluated, under whose grant, and what check it passed.

The weak answer is full conversation history. Chat logs are not attestation. They record what was said, not what was authorised — and reconstructing an action's authority basis from a transcript is an archaeology project measured in weeks, undertaken at exactly the moment nobody has weeks. If a vendor's answer to auditability is the transcript, your risk function has just been volunteered for that project without being consulted.

The distinction to hold in the meeting is between logging and attestation. Logging is a record that something happened, produced for debugging, retained on a debugging schedule. Attestation is a record produced at the moment of action, capturing what was known and what permitted it, retained on an investigation schedule. Most systems have the first. The question is whether this one has the second.

The follow-up: how long is it retained, and who can alter it? Retention that outlives a debugging window and an answer about immutability distinguish an evidence system from a log aggregator. If the record can be edited by whoever operates the system, it is not evidence in the sense that matters when someone adversarial is asking.

Question five: what does it cost per outcome at my volume?

Washed agents are priced like software — per seat, predictable, familiar to procurement. Real agents consume like infrastructure: per run, scaling with context length and step count, and rising with exactly the complexity that makes a case worth automating.

So a vendor who cannot model your cost per resolved case at production volume either has not run at production volume, or would rather you did not do the multiplication. Both answers are informative, and neither is disqualifying on its own — but both change the shape of the pilot you should be asking for.

The reason this question belongs in a filter about authenticity rather than in a pricing negotiation is that the answer is diagnostic. Per-seat pricing on a genuinely agentic system is a strong signal that the vendor has not yet met production economics, because anyone who has met them has felt the step-count multiplier and priced defensively. Confident per-seat pricing is either a system that does not consume like an agent, or a vendor about to discover that it does.

The follow-up: what happens to that number when the case is hard? Averages hide the distribution, and the expensive cases are precisely the ones you were hoping to automate. A vendor who has watched their own cost curve will know the shape of the tail.

Scoring the meeting

The score is not a verdict on the vendor. It tells you which purchase you are making, which is a more useful thing to know.

Score the meeting honestly The score is not a verdict on the vendor. It tells you which purchase you are making. 5 clean Proceed to a governed pilot You are among the ~130. The remaining risk is execution, not category. Pilot with the authority model and the evidence trail switched on from day one. 3–4 clean Find out which one failed A gap on authority or evidence is a different decision from a gap on cost modelling. One is a governance defect you would have to build around. The other is a negotiation. ≤2 clean Automation in a costume Which may still be worth buying — as automation, at automation prices, without the agent risk premium in either direction. A note for the sceptics on the board The washing is a hype symptom. It does not license the comfortable conclusion that agents themselves are vapour. The same research house forecasts 40% of enterprise applications carrying task-specific agents by 2026, up from under 5% in 2025. Buy the arrival, not the costume.

Five clean answers and you are plausibly among the genuine minority; the remaining risk is execution rather than category, and the right next step is a governed pilot with the authority model and the evidence trail switched on from day one rather than retrofitted after it works.

Two or fewer and you have found automation in a costume — which may still be worth buying, as automation, at automation prices, without the agent risk premium in either direction. That last clause matters in both directions: you should not pay the premium, and you should not apply the scrutiny either. A workflow tool does not need an authority model, and treating it as though it does wastes your risk function on a system that cannot hurt you in the ways they are checking for.

The middle band is where most real meetings land and where the filter earns its keep. Three or four clean answers is not a score to average — it matters enormously which one failed. A gap on question three or four is a governance defect you would have to build around, on someone else's roadmap. A gap on question five is a negotiation. Those are different decisions, and a single number would have hidden the difference.

Write the answers down — the meeting is an evidence artefact

One practice turns this filter from a conversation into something durable, and almost nobody does it.

Record the answers. Not a summary of the meeting — the five questions, what the vendor said to each, and the date. One page. Then treat that page as part of the procurement record rather than as a personal note that dies in an inbox.

Three things follow from that page, and each is worth more than the effort of producing it.

It converts vendor claims into commitments. A statement made in a diligence session and written into the record is a different object from the same statement made verbally. It can be attached to the contract, revisited at renewal, and pointed at when the system does not behave as described. Vendors answer question three more carefully when they can see the answer being written down, which is itself diagnostic.

It gives your risk function something to inherit. The gap I see most often is not that nobody asked good questions — it is that the person who asked them moved on, and the answers were never written anywhere the risk function could find them. Eighteen months later, when someone asks what authority model this system has, the organisation re-derives it from the vendor's marketing site because the original answers are gone. A one-page record makes the diligence a durable asset rather than a moment of individual competence.

It is the first entry in the system's evidence trail. This is the part I would press hardest. If the system turns out to matter — if it takes consequential actions and someone eventually asks who permitted this — the reconstruction starts with what the organisation understood it was buying, and on what basis. A dated record of what the vendor asserted about authority bounds and audit trails is precisely the kind of artefact that makes a later investigation tractable, and precisely the kind nobody thinks to create in advance.

There is a symmetry here worth naming. The whole filter tests whether a vendor can produce evidence about what their system did. A buyer who runs that filter and keeps no record of the answers has failed the same test one level up — asking for attestation while practising none. The page costs twenty minutes and it is the cheapest governance artefact in the entire procurement.

The strongest objection: this is gatekeeping, and the line is arbitrary

Take the counter-argument seriously, because it comes from thoughtful people and it is partly right.

It runs like this. The agent-versus-automation boundary is a definitional preference, not a fact about systems. Every useful system sits on a spectrum between fully scripted and fully autonomous, and drawing a line in the middle and calling one side real is analytically empty. Buyers should care whether a system solves their problem at an acceptable cost and risk — not which category a consultant assigns it to. And a filter like this one privileges vendors fluent in governance vocabulary over vendors who have built something that works, which selects for polish rather than substance.

Most of that is fair, and two parts of it are simply correct: the spectrum is real, and the outcome is what matters. If a re-badged workflow tool solves the problem at a good price, buy it — this filter's own scoring says so explicitly.

Where the objection breaks is on the risk posture, which is not a matter of definitional preference. The scrutiny a system attracts is driven by what it is called and what it is permitted to do, not by where a taxonomist places it. A system that can take consequential actions in your systems of record needs an authority model and an evidence trail regardless of which side of any line it sits on. So the filter is not really asking is this a real agent. It is asking: does this system take actions that need bounding, and if so, has anyone bounded them? That question is not arbitrary, and it has a factual answer.

The fluency concern is the residue worth carrying. A vendor can absolutely learn to answer all five questions well without having built any of it — the questions are public, and this piece makes them more public. Which is precisely why each has a follow-up that requires showing rather than describing: a denial demonstrated, a specific failure narrated, a retention policy stated. Vocabulary is cheap; artefacts are not.

How this lands differently by market

The filter is the same everywhere. What changes is which question fails first.

In North America, buyers generally have the model-risk and audit function to press questions three and four hard, and vendors selling into financial services have usually been forced to develop answers. The common failure is question five: sophisticated buyers, genuine capability, and nobody has modelled production-volume economics until the second invoice.

In India, the build-out is fast and the pricing pressure is severe, which selects for per-seat models and against vendors who price defensively for consumption. Question five often gets a confident answer that will not survive contact with volume. Question three is the one to press, because an authority model retrofitted later, against a vendor's roadmap, is the expensive path.

In the Gulf, much of the buying is new-build and integrated, which is the strongest position for insisting on all five before anything is committed — there is no legacy system forcing a compromise. The specific risk is that integrated stacks are bought as a whole, and the five questions get asked of the platform rather than of each agentic component inside it. Ask them component by component; the answers differ.

Where this argument is weakest

Three admissions.

First, the ~130 figure is an analyst estimate of a category with no agreed definition, produced at a point in time in a market that moves quarterly. It is directionally useful and should not be treated as a census. The argument does not depend on the number being precise — it depends on genuine capability being much rarer than claimed capability, which is not seriously disputed.

Second, the filter is biased toward detecting governance maturity, and governance maturity and agentic capability are correlated but not identical. A small vendor with genuinely novel capability and a thin authority story will score badly here. That is a real cost of the method, and the mitigation is judgement: a low score on three and four from a technically strong young vendor is a build-versus-buy signal, not a rejection.

Third, I am describing what I ask in diligence sessions, which is evidence about what surfaces useful information — not a controlled study showing that programmes using this filter outperform. Treat it as a practitioner's instrument, which is what it is.

A closing note for the sceptics on the board

There is a comfortable conclusion available from everything above, and it is wrong. The washing is a hype symptom; it does not license the view that agents themselves are vapour.

The same research house forecasting the cancellations also expects 40 percent of enterprise applications to carry task-specific agents by 2026, up from under 5 percent in 2025. Whatever one makes of any single forecast, the direction is not seriously contested by anyone building in this space.

The technology is arriving. The filter exists so that you buy the arrival, not the costume.