Picture the call, because some version of it is arriving. A customer rings her bank about a hotel charge she does not recognise — four nights, the right city, the wrong week. She did ask her assistant to book the trip. She did not ask for those dates. The analyst opens the dispute form and the form offers two options: the cardholder authorised this transaction, or she did not. Neither is true, and the analyst has to pick one anyway.
That form is the whole subject of this piece. Over 2025 and 2026 the card networks made an unusually decisive bet on agentic commerce and built, quite competently, the machinery an agent needs in order to pay: a way for a merchant to recognise that it is dealing with an agent, a way for the agent's authority to be signed and carried, a way for the transaction to route on a token provisioned for that purpose rather than on the customer's actual card number. Visa shipped Intelligent Commerce and the Trusted Agent Protocol; Mastercard shipped Agent Pay and then Agent Pay for Machines; Google's Agent Payments Protocol and Mastercard's Verifiable Intent were contributed to the FIDO Alliance in May 2026, with Visa and Mastercard chairing the payments working group that will standardise them. Authorisation is close to solved.
Dispute is untouched. The card network dispute chapters, Regulation E, and Article 4A of the Uniform Commercial Code were all drafted on an assumption so basic that nobody wrote it down: that a person pressed the button, and could therefore afterwards be asked what they meant by it. Remove the person and the assumption fails in a specific way that none of the three regimes has a remedy for. My claim is that the resulting gap is not a fraud problem and not a theft problem — those are handled — but a new and much larger middle category, where the delegation was genuine and what is contested is the execution. The gap will be litigated before it is regulated, the litigation will set the pattern that the eventual rules codify, and the institutions that are writing dispute-evidence standards now are drafting the rules everyone else will be handed.
What the networks shipped, and what they deliberately left alone
Take the accomplishment seriously first, because it is real and it is unusually fast. Visa's Trusted Agent Protocol, developed with Cloudflare and published to Visa's developer centre and a public repository, gives a merchant a cryptographic means of recognising an agent, distinguishing it from the enormous volume of undifferentiated automated traffic, and receiving structured elements describing agent intent and consumer recognition alongside the payment. Visa's own framing for why this was urgent is a cited 4,700 percent surge in AI-driven traffic to United States retail sites — a traffic figure rather than a purchase figure, and worth reading as a measure of pressure on merchant infrastructure rather than of realised agentic spend, but a real signal about what merchants are already absorbing.
Mastercard's Agent Pay is the credential-and-authorisation layer for card-funded agent transactions, sitting beneath the agent runtimes that consumers actually touch; Agent Pay for Machines, launched in June 2026, extends the same idea to high-frequency, low-value machine-to-machine payments. The Agent Payments Protocol, originally proposed by Google and donated to the FIDO Alliance, defines a vendor-neutral mandate format: a Checkout Mandate stating what the user wants to buy and under what conditions, and a Payment Mandate carrying amount, instrument and timing, each signed and each moving between an open and a closed state as the cart is assembled and finalised. Mastercard accepts mandates emitted under that format as Verifiable Intent. Stripe and Visa have both signed on. This is what convergence looks like when an industry actually wants a standard.
Now look at what did not move. Practitioner analysis of the current rulebooks finds that the Visa Core Rules have, as of April 2026, gone from indirect treatment to express language on agentic transactions — requiring identity verification in accordance with the Intelligent Commerce specifications and use of the provisioned token — while Mastercard's published Transaction Processing Rules remain silent on agentic terminology altogether, with the movement expressed through product programmes rather than through the operating rules. In both cases the change sits on the authorisation side. The dispute chapters are where they were. Liability on a card-funded agent transaction follows the ordinary tokenised allocation: where the token was validly issued and the policy honoured at authorisation, the issuer carries the fraud loss and the existing chargeback machinery runs.
That is a coherent position, and I want to be fair to it: for the fraud case it is probably the right one. If an agent's credential is stolen, or the agent itself is compromised and made to buy things nobody asked for, treating it as an ordinary unauthorised transaction is sound, because that is what it is. The problem is that fraud is not the interesting failure mode of an agent, and the networks know it. The interesting failure mode is an agent that was properly recognised, properly credentialed, acting inside a genuine delegation, that bought the wrong thing.
Regulation E has two boxes, and the agentic case is a third
Regulation E, implementing the Electronic Fund Transfer Act, is the primary federal consumer protection for disputed electronic payments in the United States, and it turns on one definition. An unauthorised electronic fund transfer is one initiated by a person other than the consumer without actual authority to initiate it, and from which the consumer receives no benefit. Everything downstream — the liability tiers, the error-resolution timetable, the provisional credit — hangs off that definition, and the definition is binary by construction. Either the authority was absent or it was not.
The binary held because, in the world it was drafted for, absence of authority was the only interesting question. Someone stole the card, someone guessed the credentials, someone at the merchant ran the transaction twice. Delegation existed — people have always given standing instructions, used brokers, let travel agents book on their behalf — but delegation was rare, bounded, and mediated by a human who could be asked afterwards what they had been told. When the customer said the travel agent booked the wrong week, there was a travel agent, with a file and a memory and a professional obligation.
An agent inverts all three properties at once. Delegation stops being rare and becomes the default mode of the channel. It stops being bounded, because a natural-language instruction has no edges — book me something nice for the conference is a scope statement no policy engine can bound and no court can construe. And the mediating party is a piece of software with no memory a reviewer can subpoena and no obligation a regulator can attach. The commentary that has looked hardest at this puts it precisely: the definition collapses the questions of authority, error and disputed authority into a single binary, and agentic commerce exposes the fragility of that collapse, because many disputed transactions will involve authorised access combined with disputed execution rather than the absence of authority altogether.
There is a subtler consequence in the timing rules, and it is the one that will surprise banks first. Regulation E's consumer liability escalates by reference to when the periodic statement was sent, on the assumption that a consumer reviewing a statement can recognise which transactions they made. An agent transacting continuously, at low value, across many merchants, breaks that assumption too. The consumer is not failing to review the statement; the consumer genuinely cannot distinguish the purchase they meant from the purchase they did not, because they never saw either one being made. A liability regime that runs a clock against a consumer's ability to notice is a regime that stops working the moment noticing becomes structurally impossible.
Article 4A allocates loss to whoever agreed the security procedure
For business payments the controlling framework is Article 4A of the Uniform Commercial Code, drafted in 1989 for commercial funds transfers, and its logic is different and in some ways more dangerous. Article 4A does not ask whether the payment order was in fact authorised. It asks whether the bank and the customer agreed a security procedure, whether that procedure was commercially reasonable, and whether the bank accepted the order in good faith and in compliance with it. If those conditions hold, the order is effective as the customer's order even when the customer did not give it, and the loss sits with the customer.
Apply that to a procurement agent and read the sentence again slowly. The security procedure in an agentic B2B flow is, functionally, the agent's own credential and the verification the network performs on it. So the framework asks a bank to assess whether the agent's authentication was commercially reasonable, and if it was, to allocate the loss of the agent's mistakes to the customer who deployed it. The agent has become the security procedure it was meant to be tested by. A framework designed to place loss on whichever party was better positioned to prevent unauthorised orders now places it on the party least able to observe what its own software did — and does so, on its face, by operation of a rule the customer signed years ago in a cash-management agreement.
I am not confident that allocation survives contact with a court, and I would not build a programme assuming it does. But the direction of the default matters enormously for who has an incentive to build evidence. Under Regulation E the consumer's exposure is capped and the institution carries the residual, so the institution has the incentive. Under Article 4A the corporate customer carries it, so the customer does. Anyone deploying a purchasing or payment-initiating agent inside a business is, under the plain text of the code, the party who will be asked to prove what the agent was told.
The mandate is an authority record wearing a payments costume
Here is the part of this story I find genuinely interesting, and it is not the gap — it is what the industry built while trying to close a different one. Read the AP2 mandate structure with a reviewer's eyes rather than a payments engineer's. A Checkout Mandate is a signed statement of what a principal authorised, expressed as scope and conditions, bound to a moment in time. Verifiable Intent turns that authorisation into portable cryptographic evidence that an issuer, a network or a merchant can verify independently of whoever is asserting it. Strip the payments vocabulary and those are the fields of an authority record: the principal, the scope as stated at grant time, the grant's boundaries, and a verification path that does not run through the party whose conduct is in question.
The Payment Mandate and the network's own authorisation message together supply something very close to a decision record — the action, the counterparty, the amount, the instrument, the effect. The provisioned token and the Trusted Agent Protocol signature supply part of a policy record: they establish that a rule about agent recognition was evaluated and returned a permit. Three of the four records that a reviewer asks for after any consequential agent action have been rebuilt, from first principles, by an industry that arrived at the problem from commerce rather than from governance and did not appear to notice the convergence. That is the strongest available evidence that these four records are not one practitioner's schema preference. They are what the structure of the problem forces.
Which brings the fourth record into focus, and it is the one every disputed execution turns on. Nothing in the agentic payments stack emits an input record: what listings the agent saw, what prices and terms were displayed to it, what it compared, what ranking it was given, what merchant-supplied text it read before deciding this hotel and these dates. When the customer says the assistant booked the wrong week and the merchant says the assistant confirmed those dates, the fact that resolves the dispute is what the assistant was shown — and that fact is held by nobody, retained by no one under any obligation, and reconstructible only by whoever happens to have logged their side of it.
This is also where the network layer and the governance layer say the same thing in different words. A provisioned token establishes that a transaction is technically capable of being processed on this account. A mandate establishes that this principal, at this time, permitted this class of purchase. A permission answers can it. An entitlement answers may it. The industry has, in the space of about eighteen months, moved from the first to the second — which is the harder half — and then stopped one record short of being able to prove it.
The merchant has the least information and the most exposure
The merchant side is where the pressure will first become visible, for a structural reason: the merchant is the only party to the transaction that has no relationship with the principal at all. The issuer knows the cardholder. The agent provider knows the user. The merchant knows a signed request arrived, verified against a protocol, from something that claimed to be acting for someone. If that transaction is later disputed, the merchant defends by producing evidence of an authorisation that happened somewhere it could not see.
The predictable response is that merchants build their own record, and I expect them to build the one nobody else is building — the input record — because it is the only evidence that helps them. A merchant that can produce a signed capture of exactly what its own surface returned to the agent, with prices, dates, terms and availability as displayed at that timestamp, has a defence. A merchant that cannot has an assertion. Notice that this asymmetry pushes the evidence layer toward the party with the least governance obligation and the strongest commercial motive, which is not where a regulator would have put it, and is exactly where a market puts things when the rules are silent.
The second-order effect deserves naming because it is already being flagged by the dispute industry: agents on both sides. Merchants deploy automated systems to detect and contest disputes; consumers acquire assistants that file and pursue them. A dispute process designed around a human reviewing a human's claim becomes a contest between two automated pipelines, adjudicated by rules that assume a person somewhere is exercising judgement. Volume rises, per-case scrutiny falls, and the outcome converges on whoever's evidence is machine-readable. That is not a fraud problem either. It is a design problem about which record wins by default.
Three markets, three clocks
United States: the fastest rails and the thinnest rulebook. This is where agentic commerce is scaling first and where the doctrinal gap is widest, because the relevant instruments — Regulation E, Article 4A, the network dispute chapters — are all mature enough that nobody is going to rewrite them quickly. The supervisory picture makes the vacuum wider rather than narrower. The revised interagency model risk guidance of 17 April 2026, Fed SR 26-2 and OCC Bulletin 2026-13, superseded SR 11-7 and expressly placed generative and agentic AI outside its scope, with separate guidance promised. That is a deferral, not an exemption: a bank's obligations for unfair or deceptive practices, for error resolution, for records, and for third-party risk attach to the action regardless of what performed it. What has been removed is the framework that would have specified the evidence — which means the first institutions to define what agentic dispute evidence looks like are, in effect, writing the baseline the separate guidance will later be drafted against.
India: the fastest adoption and the most prescriptive plumbing. India's payments infrastructure is the one place where a national-scale, real-time, low-value rail already exists at consumer scale, which makes it the most natural home for agent-initiated micropayments and also the most tightly specified. Authorisation there is a mandate question already — recurring-payment mandates, with defined limits and pre-debit notification, are an established regulatory construct rather than an industry proposal — which puts Indian institutions unusually close to the agentic answer without having framed it that way. The binding constraint is the DPDP Act: what the agent read in order to decide is personal data processed for a stated purpose, so the input record is not merely useful evidence, it is the artifact that demonstrates purpose limitation to a Data Protection Board operating on its own timetable. Of the three markets, India is where the missing record has statutory teeth soonest.
The Gulf: new-build, which is the cheap moment. Gulf institutions are deploying agentic capability inside programmes that are largely new construction rather than retrofit, and the supervisory expectations they face — the Central Bank of the UAE's February 2026 guidance note on AI adoption by licensed institutions, and comparable SAMA expectations in the Kingdom — are framed around documented governance, risk-rated inventory and board accountability for outcomes rather than around a prescriptive control list. That combination is an advantage: an institution designing its agentic payments flow now can specify the mandate and the input record as schema decisions at a cost of an afternoon, where a US or European incumbent will pay for them as archaeology. The corresponding weakness is that new-build programmes tend to record what operations needs and nothing else, and a record designed for operations is almost never the record scrutiny asks for.
What to build before the first case, and what it costs
The work divides into three pieces, and none of them is a product purchase.
- The authority side, which is largely available. Adopt the mandate structure rather than inventing one — the format is now under a standards body with the two largest networks chairing the payments group, and building a proprietary alternative in 2026 buys nothing. What is genuinely yours to decide is what your institution will treat as a valid scope: which purchase classes an agent may transact in, what ceilings apply per transaction and per period, what expiry a delegation carries, and which classes require the principal to confirm before rather than after. Those are policy decisions with legal consequences, and they are being made by default today in most institutions that have shipped anything.
- The input record, which nobody will hand you. Capture, at decision time, what the agent was working from: the candidate set, the displayed terms, the ranking, the merchant-supplied text, and the version of whatever the agent used to choose. Hold it by reference into a content-addressed store rather than copying it, so the record stays small and the material stays under its original controls. This is ordinary engineering and it is the only part of the stack that is entirely your own.
- The reconstruction path, which is the part that rots. An authenticated workflow that lets a dispute analyst — not an engineer — retrieve the mandate, the transaction, and the input record for a given purchase, in a form that can be attached to a response and read by someone hostile. Rehearse it before it is demanded. A record that has never been retrieved under time pressure is a record you do not yet know you have.
On cost: the authority side is integration work against a published specification, measured in weeks rather than quarters for an institution that already has tokenisation. The input record is storage and a capture point, and the storage is trivial — this is metadata about a purchase, not the purchase. The reconstruction path is the expensive one, because it is a real product surface with a real owner, and organisations consistently underfund it because it produces nothing until the day it produces everything. The genuine cost, as always, is temporal: every agent transaction settled before the record existed is unreconstructable and always will be, and if the first case lands inside that window the answer is a project rather than a query.
What this argument does not prove
Four limits, and the first is the largest.
This piece describes a gap, not a wave of losses. I am not aware of published data quantifying disputed agent-initiated transactions, because the volumes are still small and the disputes that exist are being absorbed into ordinary chargeback categories where they are invisible as a class. It is entirely possible that the networks are right and this resolves quietly: that agent providers indemnify commercially, that issuers absorb the middle cases as a cost of a channel they want, and that the doctrinal question never gets litigated because nobody's exposure is large enough to be worth the fee. I think that is unlikely at scale, but it is the strongest counter-argument and it deserves to be stated without hedging.
Second, my reading of the network rules is a reading. Operating rules are proprietary, revised quarterly, and what is publicly summarised lags what is in force; the Article 4A and Regulation E analysis here is architectural, aimed at where the evidence has to come from, and is not a legal opinion. Anyone making a liability decision on this should be doing it with counsel who has the current rulebooks in front of them. Third, the middle category is not new in kind. Brokers, standing instructions and travel agents have always produced disputes where the delegation was real and the execution was contested; what is new is the volume and the absence of a human intermediary who can be asked what they were told. The doctrine has handled the old version by asking that human, which is the mechanism agents remove.
Fourth, and most importantly, the input record does not adjudicate. Knowing exactly what the agent saw tells you what it had to work with; it does not tell you whether choosing those dates from that screen was reasonable. That judgement is going to have to be made by someone, against a standard that does not yet exist, and I have no idea what it will look like. What I am confident of is the ordering: you cannot build a standard of reasonable agent behaviour on top of an evidentiary layer that cannot say what the agent was looking at. The record has to exist before the standard can be written against it — which is precisely why the institutions writing records now will find their choices in the rules later.
The question worth taking into your next payments architecture review is narrower than any of this and answers most of it: if a customer disputes an agent purchase your institution processed last Tuesday, what can you produce, and how long does it take? If the honest answer is a token, an authorisation message and a shrug, the gap described here is not a forecast about the industry. It is a description of your file. If you are deciding what your agentic dispute evidence should contain before a court decides for you, that is the kind of working session I do.