Model risk & validation for agents
Validation frameworks written for systems that score, applied to systems that act.
NIST AI Risk Management Framework
Effective challenge
Ongoing monitoring
The problem
SR 11-7 was written in 2011 for models that score and predict. Your agents act. The guidance did not change — the exposure did. This lane is the one the rest of the site already covers in depth, because it is where the four-layer method lives; this page exists to place it in the stack rather than to repeat it.
Where it breaks
| Failure point | What actually happens |
|---|---|
| Inventory | Most model inventories cannot represent a system that chains tools, prompts and sub-models at runtime. |
| Validation | You validated the model. Nobody validated the agent's authority — and a well-scored model wrapped in an agent that can act outside its grant is a governance failure with excellent metrics. |
| Effective challenge | Challenge assumes the reviewer can reconstruct the decision. If the agent's context is gone, so is the evidence. |
| Monitoring | Model performance drifts slowly. Agent behaviour changes the day a provider ships a new version, with no code change and often no revalidation trigger. |
Standards and regimes named on this page were last checked against their primary sources on . Instruments move — several cited here changed inside the last year — so verify against the source before you rely on one. If you find something stale, tell me and I will correct it.
Go deeper
- The four layers, in dependency order
- Agent inventory schema
- Banking & capital markets
- Reviewer question bank
All five surfaces: capabilities.
Where does this sit in your estate?
The useful version of this conversation is specific — one workflow, one deadline, one thing you are not sure you could evidence. That is usually twenty minutes.