Updated 31 July 2026. An earlier version of this piece argued that SR 11-7 was sufficient to govern agents and that no new guidance was needed. That was defensible when written and is now wrong: the guidance itself has been superseded. The revision is set out below, along with what it changes and what it emphatically does not.
Picture the meeting that happened, in some version, inside every American bank this spring. A validation team opens the freshly issued guidance — the document their whole discipline is built on has just been replaced for the first time in fifteen years — and pages through it looking for the section that finally tells them what to do about the AI agents the business has been pushing into production. They find the section. It is one sentence long, and it says the agents are not covered.
For fifteen years the answer to “how do we govern a model in a bank” had a stable address. SR 11-7, its 2021 companion, and OCC Bulletin 2011-12 defined the units — inventory, validation, effective challenge (the requirement that someone independent can question and test the model), ongoing monitoring — and an industry of practice grew on top of them. On 17 April 2026 the Federal Reserve, the FDIC and the OCC issued revised interagency guidance on model risk management. It supersedes SR 11-7 and SR 21-8. It rescinds OCC Bulletin 2011-12, OCC 2021-19, OCC 1997-24 and the Model Risk Management booklet of the Comptroller’s Handbook. Fifteen years of scaffolding came down in a single issuance.
That alone would be worth writing about. It is not the important part. The important part is the sentence the validation team found.
The sentence that matters
Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance.
Separate guidance on AI and generative AI is promised. The revised guidance also states plainly that it “does not set forth enforceable standards or prescriptive requirements,” and it narrows the model definition to a complex quantitative method applying statistical, economic or financial theories to process input data into quantitative estimates — explicitly excluding deterministic rule-based processes and simple arithmetic.
So the framework that every governance team assumed would eventually be stretched over agentic systems has declined to be stretched. It did not fail to mention them. It named them and stepped back. Out of scope is not out of risk.
Why this is not the good news it looks like
I have already seen this read as breathing room, and that reading is a trap for a specific reason: a scope carve-out in supervisory guidance removes a framework, not an obligation.
The duties attached to the action are completely untouched. An agent that adjusts a credit limit is subject to everything that governs adjusting a credit limit. An agent that declines an application is subject to everything that governs declining an application — adverse action notice, the reasons behind it, the fair-lending posture of the decision boundary. Safety and soundness applies to the institution, not to the technique it chose. Third-party risk management does not stop applying because the model is rented. None of that lived in SR 11-7; all of it lives in the rules attached to the underlying activity, and those rules did not move an inch on 17 April.
What has been removed is the thing that would have told you what “adequate” looks like. The value of a model risk framework to a practitioner was never that it created obligations. It was that it converted a diffuse duty into a specification: here is what an inventory is, here is what validation covers, here is what effective challenge requires. Take the specification away and the duty remains, but the definition of compliance becomes a matter of institutional judgment — exercised in advance, defended afterward.
That is a harder position than being told what to build, not an easier one. It is the difference between a code and a standard of care.
Why the agencies were probably right to do it
It would be cheap to read the carve-out as regulators ducking a hard problem, and I do not think that is what happened. The old framework genuinely does not fit, and forcing the fit would have produced worse outcomes than admitting the gap.
The units of account break first. Is an agent a model, an application, or forty models composed at runtime? Most inventories cannot represent a system that assembles tools, prompts and sub-models per call, so agents get half-listed: the underlying model registered, the orchestration logged as an application, and the actual decision surface recorded nowhere.
Then the model definition itself breaks. The revised guidance defines a model as a quantitative method producing quantitative estimates. An agent that reads a case file, decides a sequence of steps, and writes to a system of record is not producing an estimate. It is taking an action. Stretching a definition written for estimates over a system that acts would have been a category error with the force of guidance behind it, and the industry would have spent three years complying with the wrong noun.
Given a choice between a bad fit and an honest gap, an honest gap is better — provided you understand that you are standing in it.
What actually fills the gap
The same four properties I would have argued for before the revision, now with a sharper reason to build them: nobody is coming to specify them for you, and the separate guidance, whenever it lands, will be written against whatever the industry has already built.
- Inventory by action class, not by artifact. “May adjust limits up to X for accounts in good standing” is a governable unit. “Frontier model, vendor-hosted” is not. This survives any definitional change in future guidance, because it is a statement about your business rather than about the technology.
- Validate the authority, not only the model. A perfectly-scored model wrapped in an agent that can act outside its grant is a governance failure with excellent metrics. Scope has to cover the authority model, the tool surface and failure containment: what can it reach, what may it do there, what happens when it is wrong.
- Enforce at the point of action. Between intent and effect, in the gateway or service layer you already operate — not as an instruction in a prompt the agent may choose to follow. If the agent can decide to bypass the check, it is not a check.
- Attest at execution. What the agent knew and from where, under which entitlement, evaluated by which policy, with what effect — written synchronously with the action. This is what turns an incident from a program-ending reconstruction project into a lookup.
One control deserves singling out because it is cheap and very frequently absent: provider version pinning with an explicit revalidation trigger on version change. Agent behavior can shift materially the day a provider ships a new version — no code change on your side, no deployment, and, in most firms, no trigger in the revalidation policy.
The honest limits of this reading
Two things I cannot tell you, and it would be dishonest to imply otherwise.
I do not know what the separate AI guidance will say or when it will arrive. It is possible it lands narrow and prescriptive, in which case some of what firms build now will need rework. My argument is not that building early guarantees alignment — it is that a firm with an authority model and an evidence trail can adapt to almost any specification, and a firm with neither cannot adapt to any of them.
And attestation for agentic systems remains genuinely immature. There is no settled schema, no interchange format, and no consensus on what a supervisor should be shown. I have views about what belongs in the record; anyone claiming a standard here is describing a product rather than a practice. I expect to be partly wrong about this in three years.
How this lands outside the United States
The export effect is real but it now runs differently, and anyone operating across regions should notice the change.
India. The binding constraint arrives through data protection rather than model risk. The DPDP Act attaches purpose limitation to personal data, which makes an agent traversing systems outside its stated purpose a live exposure independent of any AI-specific rule — and combined with the RBI’s sharpening posture on model-driven decisioning, it routes the question to the privacy function first. Indian programs frequently arrive with a well-drafted policy position and no enforcement boundary to implement it at.
The Gulf. The constraint is scale and visibility rather than examination. Sovereign-scale programs concentrate consequence, and the failure mode is political before it is operational. The compensating advantage is substantial: much of the estate is new-build, and in a new-build the four properties cost a fraction of what they cost as a retrofit.
What has changed for both is that supervisors elsewhere used to have an obvious template to reach for. That template has now formally excused itself from the agent question, which means the analogy import is weaker and the local answer has to be more original. In practice this makes an institution’s own documented reasoning — why this authority, granted by whom, enforced where, evidenced how — more load-bearing than a citation to somebody else’s framework ever was.
The position I would defend
The firms that will be comfortable in eighteen months are the ones treating the carve-out as a deadline rather than a holiday. Authority recorded at grant, actions attested at execution, revalidation tied to version change, enforcement sitting where the action crosses into the system of record.
It is not more work than the alternative. It is the same work, done while it is still cheap to change the answer — and done by the people whose practice the eventual guidance will be written to describe.
If your validation team is looking at an agent program and wondering what it is supposed to attach to now, compare notes with me. That migration is most of what I do.