I spent part of last week reading three instruments that have nothing to do with each other. Different continents, different sectors, different legal traditions, no shared drafting committee, and no evidence that any of the three drafters read the others. A central bank's draft guidance for its regulated entities. A state insurance statute. A supervisory bulletin for national banks.
They kept arriving at the same place. Not the same words — the same axis. Each one, at the moment it had to decide what kind of thing it was regulating, reached for autonomy: how much of the decision the system makes on its own, and what happens when it makes the wrong one. None of them, at any point, asked whether the model was accurate.
That convergence is more interesting than it first sounds, because it cuts against how the AI risk conversation is mostly conducted. The public argument is about capability — what models can do, how well they do it, whether they can be made to behave. The regulatory argument, where it has actually been written down, is about authority and evidence. Those are different problems with different solutions, and an organisation that has invested heavily in the first can be entirely unprepared for the second.
What India drafted
On 23–24 June 2026 the Reserve Bank of India released its Draft Guidance on Regulatory Principles for Model Risk Management, 2026, open for comment until 24 July. It reaches eleven categories of regulated entity — commercial banks, small finance banks, payments banks, local area banks, regional rural banks, urban and rural co-operative banks, NBFCs across all layers, the all-India financial institutions, asset reconstruction companies and credit information companies. That is close to the whole regulated perimeter.
The headline feature is containment. The draft contemplates that every AI system a regulated entity deploys can be overridden, suspended or deactivated through kill-switch arrangements, with human oversight over AI-driven decision-making and the ability to shut a model down immediately where it produces harmful or erroneous output. Written into a supervisory expectation, that is an architectural requirement rather than a procedural one. You either built the ability to stop the thing or you did not, and no amount of policy documentation substitutes.
The quieter feature is the consequential one. When a regulated entity assigns a risk tier to a model, the draft has it weigh not only materiality and complexity but the extent of reliance and the level of autonomy placed on the model's outputs for decision-making. Autonomy is not an aggravating footnote there. It is an input to the tiering that determines how much governance the system attracts in the first place — which means the framework prices governance by how much of the decision the machine is making.
Most model risk frameworks do not have that design, because most were written for systems that produce a number a human then uses. Tiering by autonomy assumes the opposite case exists — that some systems act — and treats it as the ordinary case rather than the exception.
This is draft guidance and the comment period closed on 24 July 2026. Draft text changes, sometimes substantially, and kill-switch language is exactly the kind of provision that attracts industry submissions. Treat it as a strong signal of supervisory intent, not as a rule in force.
What Maryland enacted
Fourteen months earlier and six thousand miles away, the Maryland General Assembly passed House Bill 820, Health Insurance – Utilization Review – Use of Artificial Intelligence. It was approved on 20 May 2025 and took effect on 1 October 2025 as Chapter 747, adding a new section 15-10B-05.1 to the state's Insurance Article.
The section opens by defining its subject, and the definition is the clause that most commentary drops.
an engineered or machine-based system that varies in its level of autonomy and that can, for explicit or implicit objectives, infer from the input it receives how to generate outputs that can influence physical or virtual environments.
Read the two clauses slowly. A system whose autonomy varies, producing outputs that change environments. That is a description of an agent, sitting in a state insurance statute, enacted in May 2025 — before most enterprise agent programmes existed.
The eleven duties
Having defined its subject that way, subsection (C) requires an entity in scope to ensure eleven things. They are worth reading in full, because the mix is unusual — some are outcome duties, some are architectural, and four of them are evidence obligations that a system either satisfies by construction or cannot satisfy at all.
- Determinations rest on the enrollee's medical or other clinical history, the individual clinical circumstances as presented by the requesting provider, or other relevant clinical information in the enrollee's record.
- Determinations are not based solely on a group dataset.
- The criteria and guidelines used for making determinations comply with the requirements of the title.
- The tool does not replace the role of a health care provider in the determination process under § 15-10B-07.
- Use of the tool does not result in unfair discrimination against enrollees prohibited by federal or state law.
- The tool is fairly and equitably applied, including in accordance with applicable regulations and guidance issued by the federal Department of Health and Human Services.
- The tool is open to inspection for audit or compliance reviews by the Commissioner.
- Written policies and procedures are included in the utilization plan filed under § 15-10B-05, including how the tool will be used and what oversight will be provided.
- The performance, use and outcomes of the tool are reviewed and revised, if necessary and at least on a quarterly basis, to maximise accuracy and reliability.
- Patient data is not used beyond its intended and stated purpose, consistent with HIPAA as applicable.
- The tool does not directly or indirectly cause harm to an enrollee.
A flat prohibition then closes the section: an artificial intelligence, algorithm, or other software tool may not deny, delay or modify health care services.
Duty seven is a standing state, not a response capability. A tool that is open to inspection for audit or compliance review is a tool whose evidence exists before anyone asks. The practical test is unkind and quick: request the inspection package internally with five business days' notice and see whether it is produced or assembled. If it was assembled, the obligation was not being met during the period under review, however good the resulting document is.
Duty eight makes your architecture a representation to a regulator. The filed utilization plan must state how the tool will be used and what oversight will be provided. That filing is a legal representation. The deployed system is a fact. Where the two diverge, the divergence is the finding — regardless of which one is better. Read the filed description aloud to an engineer who built the system and watch whether they recognise it.
Duty ten is purpose limitation, in a statute. Patient data must not be used beyond its intended and stated purpose. A flat retrieval index cannot express that constraint; it answers whatever it is asked. Satisfying this duty means a purpose boundary that refuses, evaluated at query time against the asking actor and the stated purpose — not a policy document describing a boundary nobody enforces, and not broad retrieval with the answer filtered afterwards, because in that design the data was already disclosed into the context and only the display was suppressed.
The clause that reaches the vendor
Subsection (B) sets scope, and it is broader than the headline suggests. The duties bind a carrier that uses such a tool for utilization review; a carrier that contracts with or otherwise works through an entity that uses one; and — directly, in its own paragraph — a pharmacy benefits manager or private review agent that contracts with a carrier to provide utilization review on the carrier's behalf and uses such a tool to do it.
The obligation does not stop at the payer. It travels to the party actually conducting the review. A services firm or platform running an AI-assisted utilization workflow for a health plan is not downstream of this statute; it is named by it. Contractual flow-down from the carrier is a sensible commercial arrangement and it is not the same thing as the statutory duty, which attaches directly.
The following year's Chapter 165 sharpened the reporting side. Carriers must state, in the quarterly report to the Insurance Commissioner and for each adverse decision, whether a prior authorization or step therapy protocol was involved, the type of service at issue, and whether an artificial intelligence, algorithm or other software tool was used in making the adverse decision. On examination of a pattern of adverse decisions relating to emergency department services, a carrier must produce all documents related to an adverse decision, including electronic documents in the possession of a private review agent acting on its behalf, with independent review available at the carrier's cost.
The per-decision flag has an architectural precondition that is easy to miss. Whether a tool was used in making a particular adverse decision has to be captured at decision time. It cannot be reliably reconstructed from logs afterwards, because "the tool was in the pipeline" is a materially different claim from "the tool was used in making this decision" — and only one of those is defensible in front of someone testing it. A field that must be true per decision has to be written per decision.
What the federal supervisors did
Then, on 17 April 2026, the Federal Reserve, the FDIC and the OCC issued revised interagency model risk management guidance — Fed SR 26-2 and OCC Bulletin 2026-13 — superseding SR 11-7 and SR 21-8 and rescinding the older OCC bulletins along with the model risk booklet of the Comptroller's Handbook.
SR 11-7 had been the most influential model governance document in the world. For fifteen years it was the thing every serious framework was built against, including in institutions that had never met an American regulator. Its replacement says two things that matter here. Generative and agentic AI models are novel and rapidly evolving and are not within the scope of this guidance; separate guidance is promised. And the document does not set forth enforceable standards or prescriptive requirements.
So the sequence, laid out plainly: a US state wrote autonomy into the definition in May 2025. India's central bank made autonomy a tiering input and containment an expectation in June 2026. And in between, in April 2026, the most sophisticated model risk framework in existence stepped back from exactly the systems the other two were reaching for.
What none of them asks
Strip the vocabulary and the same four demands appear everywhere a rule was actually written.
Bound what it may do. Maryland's flat prohibition — the tool may not deny, delay or modify — is a statement about permitted action, not about accuracy. RBI's kill-switch expectation is the same instinct approached from the other end: the system must have a boundary someone can enforce. A model that is right ninety-nine times in a hundred and is permitted to do something it should never have been permitted to do has not had a quality failure. It has had an authority failure, and no amount of additional accuracy addresses it.
Keep a human in the decision, not near it. Maryland requires the tool not to replace the provider's role. RBI requires human oversight of AI-driven decision-making. Both are easy to satisfy on an architecture diagram and hard to satisfy in production, because a review that is never exercised — or exercised in seconds, at volume — is a role in name. The number that settles it is median reviewer dwell time alongside the override rate, and almost nobody measures either.
Be inspectable on the regulator's timetable. A standing inspection right is a state the system is in or is not. Evidence produced under examination is not evidence of a period during which the duty applied; it is evidence of a good weekend.
Say in advance what you will do, then match it. The filed plan describes how the tool is used and what oversight is provided. The system is what it is. Divergence between the two is the finding — and it is the failure mode that grows quietly, because filings are annual and systems change weekly.
Four demands, no model quality among them. The question underneath all of them is the one that keeps arriving: who did that, on whose authority, and can you prove it?
Why a deferral is the stronger argument for building now
The obvious reading of the federal move is relief. Agentic AI is out of scope, so the burden has lifted, so the governance work can wait for the promised guidance. That reading is backwards, for two reasons.
Nothing was removed except the framework. Every obligation attached to the underlying action is untouched — safety and soundness, consumer protection and fair lending, sectoral duties, third-party risk. What went away is the document that would have specified the controls. The exposure did not move. The instruction manual did, and an exposure without an instruction manual is harder to manage, not easier.
Someone will write the sequel, against something. Separate guidance is promised. When it arrives it will be drafted by people looking at what the industry has already built, because that is how supervisory guidance has always been written. Organisations that build the evidence layer during the deferral are not merely compliant early. They are the sample the specification gets written against.
Meanwhile the states are not waiting and neither is India. A firm can be out of scope federally and squarely inside a Maryland obligation on the same workflow, in the same week, and the second fact does not care about the first.
Which is the whole of it, in the shortest form I can put it. A supervisor declining to specify controls has not declined to hold you responsible for the action — it has only declined, for now, to tell you what good looks like. Out of scope is not out of risk.
What this argument does not prove
Three limits, stated plainly, because a piece arguing for evidence should be willing to show its own.
Three instruments is a small sample. I read these three because they are recent, enacted or formally drafted, and available in primary text. I am not claiming a worldwide convergence from a sample of three, and I have not done the equivalent reading for every major regime. The pattern is real in the documents I checked. Whether it generalises is a claim I have not earned.
One of the three is a draft. The RBI text was open for comment until 24 July 2026 and its final form is unknown. The Maryland material is different in kind — both chapters are enacted, and everything quoted here is taken from the chapter text rather than from a summary of it.
Nobody has enforced any of this yet. A statutory inspection right that has never been exercised tells you what a regulator may do, not what it will do or how hard. The honest position is that the obligations exist and the enforcement posture is unknown.
What survives all three caveats is narrow and sturdy: where a regulator has recently written a rule for systems that act, the rule has been about containment, authority and evidence. Not one of them asked about accuracy.
Two claims about HB 820 circulate widely in secondary commentary and do not appear in the enacted chapter. The first is that the final determination must be made by a physician of the same specialty; the Act says the tool must not replace the role of a health care provider in the determination process under § 15-10B-07, and cross-references that section. The second is that the quarterly review covers effectiveness, accuracy and fairness; the enacted wording is performance, use and outcomes, reviewed and revised to maximise accuracy and reliability. Both summaries are close enough to sound right, which is why they spread. Quote the chapter.
The four things to have ready
If you run anything agentic in a regulated function, these four sit at the intersection of what all three instruments assume you can produce. None of them requires the promised federal guidance to exist first.
- An action boundary you can demonstrate technically — not a policy stating what the system will not do, but a test showing it cannot. If the only thing preventing an automated actor from issuing a final decision is process, then process is the control, and process is not what any of these instruments contemplates.
- A per-decision record, written at decision time — including whether an automated tool was used in making that specific decision, captured as it happens rather than inferred later from infrastructure logs.
- Oversight evidence rather than an oversight design — what the reviewer saw, when, how long they had, and how often they departed from the recommendation. An override rate of zero is itself the finding.
- An inspection package that already exists — testable today by requesting it internally with five business days' notice and observing whether it is produced or assembled.
The full mapping — all eleven duties, the evidence each requires, and a test you can run for each, together with the reporting and production obligations and the architectural precondition each one carries — is published as an open artifact. It is free, requires no email, and is built entirely from the enacted chapter text.
The systems being built right now will still be running when the promised guidance arrives. The ones that can answer for themselves will survive an audit. The ones that cannot will be explained, at length, by somebody who was not in the room when the architecture was chosen.