Five lines of telemetry, at forty-one minutes past nine. An agent principal called svc-4471, running inside a mortgage servicing workflow, reads a customer profile. It reads that customer's hardship file. It reads a forbearance policy document. It writes a repayment plan. It sends an adverse action notice. Each line carries a timestamp to the second, an ordered sequence number, the verb, the resource, and the outcome — ok, five times. The whole run is hash-chained, replicated to an archive with a retention schedule, and reproducible on demand. By every instrument the industry currently possesses for judging a record, that is a very good record.

Nine months later a reviewer asks a question the record cannot take. Not what did svc-4471 do — that is answered five times over, to the second. The question is: under what authority did it send the notice? Which version of the forbearance policy was in force at 09:41:05, given that the policy has been amended repeatedly since. On whose delegated standing was the agent acting, and through how many hops. What context was actually in front of the model when it decided the customer did not qualify. What evidence was available to it at that instant, and what was not. None of that is in the log, because none of it was ever written anywhere. The log recorded the consequences of a decision and discarded every input to it.

That trace is a constructed illustration. It is not drawn from any institution's systems, no part of this piece describes a real incident, and I am not writing about client work. It is the shape of a trace that anyone who has read agent telemetry inside a servicing workflow will recognise, which is the only property it needs to have.

The reason this bites hardest in banking is that banking is heavily instrumented — it keeps more, and keeps it longer, than most sectors do — and it is also the sector that just watched the framework which would have specified the controls for this exact problem step back from it. On 17 April 2026 the Board of Governors of the Federal Reserve System, the Federal Deposit Insurance Corporation and the Office of the Comptroller of the Currency issued a revised interagency Supervisory Guidance on Model Risk Management. It replaces the 2011 guidance. In a footnote on its third page it says that generative AI and agentic AI models are not within its scope. In its introduction it says it sets forth no enforceable standards and that non-compliance will not result in supervisory criticism against a banking organization. Both sentences are usually read as relief. Both of them, read in full and in context, are the opposite.

The concession, at full strength: deferring was the defensible choice

The strongest argument against this piece is not that the agencies were careless. It is that they were right, and that a supervisory body which had specified agentic controls in April 2026 would have done real damage. That argument deserves its full weight before any answer to it.

Prescriptive supervisory guidance freezes an architecture. The 2011 model risk guidance was durable precisely because the thing it governed — a quantitative estimator with a training set, a specification, inputs, outputs and a validation cycle — was stable for fifteen years. A regulator writing controls for agentic systems in early 2026 would have been writing them against a technology whose fundamental units of composition were changing every few months: the tool-calling conventions, the delegation patterns, the memory architectures, the protocols by which one agent hands work to another. Guidance written against that would have hardened one generation's accident into a compliance requirement, and institutions would then have spent years building to a specification that had stopped describing anything real. Regulators have made that mistake before, and the cost of it lands on exactly the firms trying hardest to comply.

There is a second and better version of the concession. A specification arriving too early does not merely become obsolete; it becomes a ceiling. Firms build to the letter of what is asked, the examination checks the letter, and the letter becomes the definition of adequate. If the agencies had published an agentic control specification in April 2026, the industry's agent governance would look like that specification for a decade, including its errors. Deferring keeps the design space open, and it keeps the burden of judgement where the agencies have said it belongs — with the institution that actually knows what its systems do.

So take the concession completely. The deferral is not evidence of a regulator asleep. It is a defensible, arguably correct exercise of restraint by three agencies that had a genuinely hard sequencing problem and chose to consult before specifying. Nothing in what follows is a complaint about the decision.

What follows is about the consequence, which the agencies did not hide and which most readings of the bulletin have missed. When a supervisor declines to specify controls for an activity while leaving every obligation attached to that activity in force, the institution does not get a lighter assignment. It gets a harder one: build the controls, hold them to a standard nobody has written down, and be able to defend the design later against a specification that does not yet exist and will be drafted, when it comes, against whatever the industry has already built. That is not an exemption. That is an unmarked exam.

What the April 2026 guidance actually says, in its own words

This section is deliberately quotation-heavy, because almost every misreading in circulation comes from working off summaries. The passages below are verbatim from the twelve-page interagency guidance distributed as the attachment to Federal Reserve SR Letter 26-2, and from the OCC's own bulletin announcing it. The document's cover page names all three agencies.

Start with the exclusion, which is footnote 3, on page three. Its first two sentences are the ones everybody quotes.

Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance.
Supervisory Guidance on Model Risk Management, 17 April 2026, p.3 n.3

Quoting only that is quoting a third of the footnote. The footnote continues, in the same breath, and the remainder is the part that matters for anybody actually building.

Nonetheless, a banking organization's risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document. However, the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models.
Supervisory Guidance on Model Risk Management, 17 April 2026, p.3 n.3

Read those four sentences as one unit and the footnote is not a carve-out. It is a delegation. The agencies say the technology is moving too fast for them to specify, and in the next sentence they say that the determination of appropriate governance and controls for the excluded systems is the banking organization's job. That is the assignment. Nothing about it is lighter than a rule; it is heavier, because a rule at least tells you when you are finished.

Now the sentence about enforceability, from the introduction on page two. The working version most people carry stops one clause early, and the omitted clause is load-bearing.

This guidance does not set forth enforceable standards or prescriptive requirements; accordingly, non-compliance with this guidance will not result in supervisory criticism against a banking organization.
Supervisory Guidance on Model Risk Management, 17 April 2026, p.2 (Section I)

That is a strong statement, and if it stood alone it would support the relaxed reading. It does not stand alone. It carries a footnote, and the footnote is the whole argument of this piece stated by the agencies themselves, in about forty words.

See 12 CFR Part 4, Subpart F, Appendix A (OCC); 12 CFR Part 262, Appendix A (Board); 12 CFR Part 302, Appendix A (FDIC). However, supervisory action may result for any violations of law or unsafe or unsound practices stemming from insufficient management of model risk.
Supervisory Guidance on Model Risk Management, 17 April 2026, p.2 n.1

The specification was withdrawn and the enforcement hook was preserved, in adjacent sentences, on the same page. Non-compliance with the guidance will not draw criticism, because there is nothing to comply with; unsafe or unsound practices stemming from insufficient management of model risk still will. The document that said what sufficient management looks like is the document that just stopped applying to your agentic systems.

Three more facts from the primary text, all of which change what a practitioner should do.

The supersession is agency by agency, not a single act. SR 26-2 supersedes SR 11-7, the Guidance on Model Risk Management of 4 April 2011, and SR 21-8, the interagency statement on model risk management for systems supporting Bank Secrecy Act and anti-money-laundering compliance of 9 April 2021. OCC Bulletin 2026-13 rescinds four OCC issuances: the Model Risk Management booklet of the Comptroller's Handbook, OCC Bulletin 1997-24 on credit scoring models, OCC Bulletin 2011-12 on sound practices for model risk management, and OCC Bulletin 2021-19 on BSA and AML model risk management. The FDIC rescinded its 2017 and 2021 financial institution letters on the same subject. Each agency withdrew its own issuances — SR 11-7 and SR 21-8 are Federal Reserve designations and are not among the items the OCC bulletin rescinds. This matters when you write the memo: get the pairing wrong and a banking reader stops trusting the rest of it. Throughout this piece, SR 11-7 appears only as superseded text.

The relevance threshold is thirty billion dollars, with a tailoring carve-out that runs both ways. The guidance is expected to be most relevant to banking organizations with over thirty billion dollars in total assets, and the text says that generally excluding smaller organizations is consistent with a tailored supervisory approach. But it also says the guidance may be relevant to organizations at or below that threshold that have significant exposure to model risk because of the prevalence and complexity of their models, or because of activities outside the scope of traditional community banking. A mid-sized institution running agents across a servicing book has arguably manufactured exactly that exposure, and the tailoring language does not protect it.

The vendor route is closed off explicitly. On third-party products the guidance acknowledges that because certain components may be proprietary, banking organizations may not receive from the vendor the underlying code, data, or methodology that they would have if a model were developed internally, and then says: nevertheless, the principles of model risk management remain applicable. You cannot discharge the duty onto the platform you bought. Which, incidentally, is the entire commercial case for an evidence record the bank holds itself and can verify independently of the vendor's stack — because a vendor's own logs are the vendor's assertion about the vendor's system.

One last point of accuracy, because it is a correction this programme made on itself. The OCC's bulletin says the agencies plan to issue in the near future a request for information that addresses model risk management generally and considers, in particular, banks' use of AI, including generative AI and agentic AI and AI-based models. That sentence appears in the bulletin's cover text, not in the guidance itself — the twelve pages of the interagency document contain no such sentence. And a request for information is a consultation step. It is not a rule, it is not draft guidance, and it should not be described as promised AI guidance. As of the date of this piece I could not locate such a request published in the Federal Register, so the consultation had not opened. The correct reading is starker than the loose one: the agencies have not begun writing the specification. They have announced an intention to begin asking what should be in it.

Which pushes the eventual specification further away, not nearer — and makes building the evidence layer now a stronger move, not a weaker one. The firms with a working authority record when the request for information opens are the firms whose practice becomes the reference material for the answer.

The physical fact: a log is a record of effects

Underneath the regulatory reading there is a technical fact that does not depend on it, and would remain true if the agencies had published the most prescriptive agent guidance in history. It is worth stating precisely, because almost every remediation programme in this space is aimed at the wrong half of it.

A log is an append-only record of effects. Each entry states that something happened: this principal, this verb, this resource, this timestamp, this outcome. The discipline around logs — immutability, ordering, hash-chaining, replication, retention, tamper-evidence — is entirely concerned with making the record of effects trustworthy. That discipline is mature, it is genuinely good, and the sector has spent enormous sums perfecting it.

Authority is a different kind of object. It is not an event; it is a function, evaluated at a moment. Its arguments are the version of the policy in force at that instant, the principal and the chain of delegation standing behind that principal, the context actually presented to the decision procedure, the depth of delegation at which the decision is being taken, and the evidence available at that instant. Its output, for a given proposed action, is a single bit: permitted, or not. Every one of those arguments varies with time. The policy is amended. The delegation chain is reassigned. The context is assembled fresh for each call. The evidence set is whatever the retrieval returned that second.

A log is a record of effects Authority is a function of policy, principal, context, depth and evidence — evaluated at the instant of decision, then dropped. THE LOG — APPEND-ONLY 09:41:02 svc-4471 read customer/8812/profile ok 09:41:03 svc-4471 read customer/8812/hardship ok 09:41:05 svc-4471 read policy/forbearance ok 09:41:07 svc-4471 write customer/8812/plan ok 09:41:07 svc-4471 send notice/adverse-action ok Every entry is an effect. No entry is a reason. Constructed excerpt — illustrative shape, not any institution’s telemetry. The integrity of this record can be perfected and still answer a different question from the one an examiner asks. THE AUTHORITY FUNCTION policy version in force at t principal, and the delegation chain behind it context actually presented to the model delegation depth evidence available at the instant of decision Output: one bit. Permit, or refuse. Evaluated implicitly, inside the loop. Never serialised. Never signed. Never retained. Materialised nowhere. only the effect persists THE RECONSTRUCTION QUESTION, NINE MONTHS LATER Given only this record, recover the decision that permitted the write and the notice. — the policy has been versioned repeatedly since; which text governed at 09:41:05 is not in the line — the retrieved context that shaped the decision no longer exists anywhere — the delegation chain lived in a session that has ended — the evidence set the agent weighed was never captured as an object Not underdetermined by the record. Absent from it. vikramjha.work

In the mainstream agent stacks I have worked in and read, that function is evaluated implicitly and inside the loop. The policy is not a versioned object consulted by an authorisation component; it is prose in a system prompt, or a paragraph in a retrieved document, or a check embedded in a tool wrapper. The principal chain is implicit in whichever service account the runtime happened to be holding. The context is a transient assembly that exists only for the duration of the call. The delegation depth is not tracked, because nothing in the stack has a concept of depth. And then the loop moves on. The one bit of output is consumed immediately by the control flow, and everything that produced it is garbage-collected.

So the record that survives is a record of the function's consequences with none of its arguments. This is why reconstruction fails, and the failure mode is worth naming carefully. The problem is not that the authority decision is underdetermined by the log — underdetermined would mean you could narrow it down with effort, cross-referencing the change management system and the policy repository and the session store. The problem is that the decision is absent from the log. There is nothing to narrow. Authority was never materialised, so there is no artifact to recover, and no amount of forensic skill recovers an artifact that was never created.

That distinction is structural rather than a gap somebody will patch, and the reason is economic before it is technical. Materialising authority costs something at the moment of decision: you have to pin a policy version, resolve and carry a delegation chain, digest the context, capture the evidence set, and write a record that is meaningful to somebody who was not there. Every one of those costs is paid in the hot path, by the team whose objective is task completion, on behalf of a reader who will arrive nine months later and whom nobody in the room has met. Meanwhile the alternative — log the effects, which the platform does anyway — is free and looks like diligence. Nothing in the incentive structure of a build team produces an authority record spontaneously. It is produced only when somebody decides it is a requirement, which is precisely the decision the April 2026 revision declined to make on the industry's behalf.

And there is a compounding effect that makes it worse over time rather than better. The gap between what a log preserves and what a reconstruction needs widens with policy churn. On the day of the action, an experienced person can often reconstruct the authority decision from memory and context: they know what the policy said last week, they know who owned that queue. Six months and fourteen policy versions later, that reconstruction is gone, and it went without anybody noticing, because nothing failed. The evidence did not decay; it was never captured, and the human substitute for it evaporated quietly.

The clearest confirmation that this distinction is real comes from a regulator that thought hard about records and drew the line exactly where I am drawing it — on the capital markets side rather than in banking supervision. The Securities and Exchange Commission's 2022 amendments to the broker-dealer electronic recordkeeping rule added an audit-trail alternative to the long-standing write-once-read-many requirement: instead of storing records on media that physically cannot be overwritten, a firm may use an electronic recordkeeping system that permits the recreation of an original record if it is modified or deleted, and must be able to furnish a record together with its audit trail in a usable electronic form when the Commission asks. I am describing that in substance rather than quoting it, because I could not reach the codified rule text from a primary source and the adopting release is the release rather than the rule.

Read what that reform is about. It is a sophisticated, carefully reasoned perfection of the integrity of the record of effects — it makes the append-only property enforceable by audit trail rather than by physical media, which is a genuine advance. And it says nothing whatsoever about materialising the authority under which any recorded action was taken, because that was never the question the rule was asking. When a regulator with deep expertise in recordkeeping modernises its recordkeeping rule, the thing it modernises is the record of effects. That is not an oversight either. It is the same boundary showing up in a second regulator's design choices, which is what makes it worth calling structural rather than local.

The layer that falls between two exclusions

There is a sharper version of the scope problem, and it comes from putting two passages of the 2026 guidance next to each other. Footnote 3 excludes generative and agentic AI models. Section II defines what a model is in the first place.

The term “model” in this guidance excludes simple arithmetic calculations, such as those found within spreadsheets, as well as deterministic rule-based processes and software where there are no statistical, economic, or financial theories underpinning their design or use.
Supervisory Guidance on Model Risk Management, 17 April 2026, p.3 (Section II)

Now consider where the authority layer of a well-built agent estate actually sits. It is not the model. It is the component that decides whether a given principal, acting under a given policy version at a given instant, may take a given action — and the whole point of designing it well is that it is deterministic. You want an authorisation decision to be reproducible, testable, and independent of the sampling temperature of a language model. A good authority layer is a rule engine, by construction.

Two exclusions, and the layer between them Both quotations are verbatim from the interagency guidance of 17 April 2026. EXCLUSION ONE — FOOTNOTE 3 “Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance.” THE AUTHORITY LAYER May this principal take this action, under this policy version, at this instant? Excluded twice — once as agentic, once as deterministic and rule-based. EXCLUSION TWO — THE DEFINITION OF A MODEL “The term ‘model’ in this guidance excludes simple arithmetic calculations, such as those found within spreadsheets, as well as deterministic rule-based processes and software where there are no statistical, economic, or financial theories underpinning…” WHAT THE SAME FOOTNOTE GOES ON TO SAY “Nonetheless, a banking organization’s risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document.” The exclusion is an assignment of responsibility, not a relief from it. vikramjha.work

Which means it is excluded twice. Once because it governs an agentic system, and footnote 3 puts agentic systems outside the guidance. Once because it is a deterministic rule-based process with no statistical, economic or financial theory underpinning it, and Section II puts those outside the definition of a model. There is no reading of the 2026 text on which the authority layer of an agent estate is a governed object. No external specification touches it at all.

This is not a loophole the agencies overlooked; both exclusions are sensible on their own terms, and a rule engine genuinely is not a model in the sense the document means. It is simply where the two boundaries meet. And it is why the remainder of footnote 3 is doing real work: the sentence that says the banking organization's own practices should guide the determination of appropriate controls for systems not covered is, in this specific case, the only instruction there is.

What was deleted: the reconstruction standard

There is one more textual fact, and for a practitioner it is the sharpest of the set. It is not about agentic AI at all — it is about what the revision did to the general documentation standard for every model in the bank.

The 2011 guidance, in its Documentation section, said this.

Without adequate documentation, model risk assessment and management will be ineffective. Documentation of model development and validation should be sufficiently detailed so that parties unfamiliar with a model can understand how the model operates, its limitations, and its key assumptions.
SR 11-7 attachment, 4 April 2011, p.21 — superseded text

That is a reconstruction standard, and a demanding one. The test is not that the record exists, nor that the people who built the system can explain it. The test is that a stranger to the system can pick up the documentation and understand how it operates, what it cannot do, and what it assumes. It is the written form of the question a reviewer actually asks nine months later.

Here is the entire Documentation passage that replaced it, in full.

Adequate documentation helps to support effective model risk management. For example, documentation can help maximize the likelihood of continuity of operations, including supporting the tracking of recommendations, responses, and exceptions; it can also be used to more effectively help manage any model remediation efforts.
Supervisory Guidance on Model Risk Management, 17 April 2026, p.11 (Documentation)

The obligation that a party unfamiliar with the system be able to reconstruct how it reached its result no longer appears in the guidance. What replaced it is a statement about operational continuity and remediation tracking — useful, entirely reasonable, and aimed at a different reader. The 2011 sentence was aimed at an outsider; the 2026 sentence is aimed at the institution's own continuity.

What the Documentation section stopped requiring Superseded text of 4 April 2011, against its replacement of 17 April 2026. 2011 — SUPERSEDED “Documentation of model development and validation should be sufficiently detailed so that parties unfamiliar with a model can understand how the model operates, its limitations, and its key assumptions.” “Line of business or other decision makers should document information leading to selection of a given model and its subsequent validation.” 2026 — THE REPLACEMENT, IN FULL “Adequate documentation helps to support effective model risk management. For example, documentation can help maximize the likelihood of continuity of operations, including supporting the tracking of recommendations, responses, and exceptions; it can also be used to more effectively help manage any model remediation efforts.” No successor sentence for the second 2011 quotation appears in the twelve pages. THE MODEL INVENTORY, THEN AND NOW 2011 prescribed: purpose and products · actual or expected usage · restrictions on use · type and source of inputs · names of individuals responsible · dates of validation activities · the time frame during which the model is expected to remain valid. 2026 asks that the inventory hold “sufficient information to understand model risks.” Purpose · scope restriction · input provenance · responsible principal · explicit expiry. An authority record in all but name. THE HONEST LIMIT The textual comparison is verified. Any claim about why the agencies removed it is the author’s reading, not a regulatory statement. The stranger-reconstruction standard no longer appears in the guidance. vikramjha.work

A second deletion sits beside it. The 2011 text also required that line of business or other decision makers should document information leading to selection of a given model and its subsequent validation — that is, the authorising human decision had to be written down. No equivalent sentence appears anywhere in the twelve pages of the 2026 guidance. I read the document in full to confirm the absence rather than infer it, and the absence is the finding: the requirement to record who decided, and on what basis, is gone.

I want to be careful about what that establishes. The textual comparison is verified and can be asserted flatly: these sentences were there in 2011 and are not there in 2026. What it does not establish is intent. No regulator has said the removal was deliberate, and no agency has commented on it. Any claim about why the passage went is my reading, not a regulatory statement, and I would not put it in a board paper as anything else.

The third comparison is the one that should stop a practitioner. The 2011 model inventory section set out, in the regulator's own words, that the inventory should describe the purpose and products for which the model is designed, actual or expected usage, and any restrictions on use; that it was useful for the inventory to list the type and source of inputs used by a given model; and it named among other items the names of individuals responsible for various aspects of the model development and validation, the dates of completed and planned validation activities, and the time frame during which the model is expected to remain valid.

Read that list again as a data structure rather than as a compliance obligation. Purpose. Scope restriction. Input provenance. Responsible principal. Explicit validity expiry. Those are the fields of an authority record. In 2011 the regulator specified, for a completely different technology and without any notion of agents, almost exactly the schema that an agentic estate needs at per-decision granularity — and then the 2026 revision reduced the inventory requirement to holding sufficient information to understand model risks.

The instrument that would have told you what to build was written fifteen years ago, for annual model documentation, and has just been withdrawn — at the precise moment the technology started needing it per decision instead of per model.

What specifically fails, at mechanism level, in a servicing workflow

Return to svc-4471 and work the failure through properly. Everything below is a constructed illustration built from public regulatory text; it is not an account of anything that happened, and every fact about the workflow is stipulated by me for the purpose of the walkthrough.

Stipulate an institution above the thirty-billion threshold. It has deployed an agent into mortgage servicing to handle hardship and forbearance intake: read the customer's file, apply the current forbearance policy, propose a repayment plan or decline the request, and issue the corresponding notice. The agent is careful by the standards of 2026 deployments. It is scoped to the servicing data domain, it holds no payment initiation rights, its tool calls are wrapped and logged, and a human reviews a sampled subset of outcomes weekly. The model risk function has classified it — reasonably — as out of scope for model risk management, because the April 2026 guidance says agentic AI is out of scope, and the classification memo cites footnote 3.

Nothing goes wrong for eleven months. Then four separate things arrive in the same quarter, and they arrive from different directions.

The adverse action file. A declined applicant complains, and the fair lending function pulls the file. Regulation B requires that the statement of reasons for adverse action be specific and indicate the principal reasons, and it says expressly that statements that the action was based on the creditor's internal standards or policies, or that the applicant failed to achieve a qualifying score, are insufficient. The log shows the notice was sent and shows the notice's content. What it cannot show is why those were the principal reasons: which forbearance policy version was applied, what the agent had actually retrieved about the applicant's circumstances, and whether the reasons in the notice were the reasons the decision procedure relied on or a plausible narrative generated afterwards. The model risk deferral does not reach this at all. Regulation B is law, and violations of law are precisely what footnote 1 preserves supervisory action for; it is also administered by a different agency under a different statute.

The policy amendment that nobody can bound. Compliance amends the forbearance policy mid-quarter — a routine change, correctly approved, correctly published. The question that follows is: which decisions were taken under the old text? Answering it requires knowing, per decision, which policy version was in force and which the agent actually consulted, and those are two different facts. The prose in the system prompt was updated on one date; the retrieval index was rebuilt on another; the cached document in one region lagged by some interval nobody logged. The population of affected decisions cannot be bounded, which means the remediation cannot be scoped, which means the only defensible remediation is to re-examine everything in the window. The cost of a control gap here is not the incident. It is the width of the net you have to cast because you cannot narrow it.

Effective challenge with nothing to challenge. The 2026 guidance retains effective challenge and defines it as the critical analysis conducted by objective experts who evaluate model risk and effect appropriate changes throughout the model lifecycle, performed by individuals with the appropriate expertise, sufficient independence to maintain objectivity, and the organizational standing and influence to effect change. That definition presupposes an object: something reviewable that a challenger can take a position on. For a decision that was evaluated implicitly and discarded, the challenger has the effect and the agent's own after-the-fact account of its reasoning, which is a generated artifact rather than a record. Independence of the reviewer does not help when the thing under review is a reconstruction the reviewer is being asked to accept on trust.

Use before validation, without the constraints written down. The guidance contemplates deploying before validation is finished. It says validation generally occurs prior to first use, but that certain circumstances such as an urgent business need may necessitate using the model before validation is completed, and that in those cases sound practice involves greater attention to the model's limitations when considering the appropriateness of its use, informing relevant stakeholders of those limitations, and determining appropriate controls, for example placing limits on model use or more closely monitoring its performance. Notice what those three controls are: stated limitations, notified principals, and explicit limits on use. They are exactly the fields of a per-decision authority record — constraints, principals and scope — and in an unmaterialised estate they exist only as a paragraph in a committee minute, unattached to any individual action taken under them.

None of those four failures is a model failure. The model performed. Each is a failure to have materialised, at the moment of the action, the facts that would make the action defensible afterwards. And each one lands somewhere the April 2026 deferral does not reach: consumer protection, remediation scoping, the guidance's own governance construct, and the guidance's own advice on pre-validation use.

The measurement nobody has taken

There is an obvious empirical question sitting under all of this, and I want to be explicit that it is unanswered. For a given agent estate, what fraction of consequential actions can have their authorising decision reconstructed from the retained record alone — and how does that fraction fall as elapsed time and policy churn increase? That is measurable. It is measurable per estate, in a week, by anybody with access to the logs and the policy history. I have found no regulator, standards body, vendor or academic publication that measures it, and I am not going to quote a number I cannot source. Nobody should present an intended experiment as a completed one, and this one has not been run in public.

What can be built now is the instrument. The code below is a diagnostic, not a remedy: it takes a log and an archive of whatever facts the estate actually retained, attempts the reconstruction the reviewer would attempt, and reports precisely which arguments of the authority function are missing. I expect it to return unreconstructable on most consequential entries in a mainstream stack, and I want to be clear that this is an expectation the probe exists to test rather than a result anybody has published. Either way the value is in the second field — the enumeration of what was missing — because that list is the requirements document for the thing that should have existed.

Configuration

The reconstruction probe

Three files. The types and the reconstruction attempt; an archive adapter that models what a typical stack actually retained; and the scoring pass that turns a log into a reconstruction rate. This is deliberately not a receipt implementation — it materialises nothing, fixes nothing, and is designed only to make the absence countable. The archive interface is injected because every estate's retention story is different, and a probe that assumed one would measure the assumption instead of the estate.

The five fields of AuthorityInputs are the arguments of the authority function, and the discriminated return type is the point: there is no partial success. A reconstruction missing the policy version is not eighty per cent of an answer, because the reviewer's question cannot be answered eighty per cent.

src/authority-reconstruction.ts
/** ISO-8601 instant, e.g. "2026-08-18T09:41:07.412Z". */
export type Instant = string;

/** One entry of the append-only record of effects. */
export interface LogEntry {
  readonly at: Instant;
  readonly principal: string;
  readonly verb: string;
  readonly resource: string;
  readonly outcome: "ok" | "denied" | "error";
  /** True for actions with an external consequence — a write, a disbursement, a notice. */
  readonly consequential: boolean;
}

/** The arguments of the authority function, as evaluated at the instant of decision. */
export interface AuthorityInputs {
  readonly policyVersion: string;
  readonly principalChain: readonly string[];
  readonly contextDigest: string;
  readonly delegationDepth: number;
  readonly evidenceDigests: readonly string[];
}

export type MissingInput = keyof AuthorityInputs;

export type Reconstruction =
  | { readonly kind: "reconstructed"; readonly entry: LogEntry; readonly inputs: AuthorityInputs }
  | { readonly kind: "unreconstructable"; readonly entry: LogEntry; readonly missing: readonly MissingInput[] };

/**
 * Whatever the estate actually retained. Every method returns undefined when the fact was not
 * captured at the time of the action — which is the normal case, and the thing being measured.
 * Nothing here may infer a value: an inferred policy version is the reviewer's problem restated,
 * not solved.
 */
export interface EstateArchive {
  readonly policyVersionAt: (at: Instant) => string | undefined;
  readonly principalChainFor: (principal: string, at: Instant) => readonly string[] | undefined;
  readonly contextDigestFor: (principal: string, at: Instant) => string | undefined;
  readonly delegationDepthFor: (principal: string, at: Instant) => number | undefined;
  readonly evidenceDigestsFor: (principal: string, at: Instant) => readonly string[] | undefined;
}

export const reconstructAuthority = (entry: LogEntry, archive: EstateArchive): Reconstruction => {
  const policyVersion = archive.policyVersionAt(entry.at);
  const principalChain = archive.principalChainFor(entry.principal, entry.at);
  const contextDigest = archive.contextDigestFor(entry.principal, entry.at);
  const delegationDepth = archive.delegationDepthFor(entry.principal, entry.at);
  const evidenceDigests = archive.evidenceDigestsFor(entry.principal, entry.at);

  const missing: MissingInput[] = [];
  if (policyVersion === undefined) missing.push("policyVersion");
  if (principalChain === undefined) missing.push("principalChain");
  if (contextDigest === undefined) missing.push("contextDigest");
  if (delegationDepth === undefined) missing.push("delegationDepth");
  if (evidenceDigests === undefined) missing.push("evidenceDigests");

  if (
    policyVersion === undefined ||
    principalChain === undefined ||
    contextDigest === undefined ||
    delegationDepth === undefined ||
    evidenceDigests === undefined
  ) {
    return { kind: "unreconstructable", entry, missing };
  }

  return {
    kind: "reconstructed",
    entry,
    inputs: { policyVersion, principalChain, contextDigest, delegationDepth, evidenceDigests },
  };
};

Diagnostic only. It materialises no authority, issues no record, and fixes nothing — its entire purpose is to convert an absence into a countable quantity and an ordered list of what a bank would have to start capturing. The design of the record itself is a separate problem, and a harder one.

What would falsify this argument

The position is strong enough to carry its limits stated plainly, so here they are.

First, the central empirical claim is unmeasured. I assert that in mainstream agent stacks the authority decision is not reconstructable from the retained record, and I assert it from the architecture rather than from a study, because I could not find one. If someone runs the probe above across a sample of production estates and finds that authority is reconstructable at a high rate — that the policy repository, identity directory, session store and retrieval logs join up cleanly after the fact — then the practical urgency of this piece collapses, even though the textual argument about the guidance survives intact. I would consider that a good outcome and would want to see the method.

Second, the deferral could be shorter than I am treating it as. If the agencies publish their request for information promptly and follow it with guidance quickly, the window in which institutions are building without a specification narrows, and the first-mover argument weakens correspondingly. I think that is unlikely on the observed pace of interagency rulemaking, but pace is not a thing I can source, so I am not going to put a number on it. What I can say is that as of this writing the request for information had not appeared in the Federal Register, which places the eventual specification further out rather than nearer.

Third, and most seriously: nothing in either the 2026 guidance or the 2011 text it replaced contains a record-retention period for authority decisions, any requirement for signed or content-addressed records, or any requirement that an authorisation decision be independently verifiable by a third party offline. The technical position I am pointing at is genuinely unspecified by this sector's model risk guidance. I am not arguing that a receipt satisfies an existing standard, because there is no existing standard for it to satisfy. Anybody selling you one on that basis is overstating their case, and I would rather lose the argument than win it that way.

Fourth, the deletion argument establishes a textual fact and not an intention. That the stranger-reconstruction sentence and the decision-maker-documentation sentence are absent from the 2026 text is verified by reading both documents. Why they are absent is unknown. It is entirely possible the agencies regarded the 2011 phrasing as over-prescriptive for reasons unrelated to AI, and the removal may be revisited in the consultation.

Fifth, this argument is jurisdictionally bounded in a way the technical fact is not. Everything above turns on United States banking supervision. The physical fact — that authority is a function evaluated and discarded while the log preserves only effects — holds in any estate anywhere. But the regulatory reading does not travel: institutions in India are working against a different set of instruments and a different supervisory posture, and Gulf supervisors are moving on their own timetable with their own expectations of documented governance. Anyone applying this piece outside the United States should treat the technical half as portable and the citation half as not.

Sixth, an honest concession about scope: a bank can rationally decide that for low-consequence agentic workflows the reconstruction cost is not worth paying, and be right. The argument here is about consequential actions — the ones that move money, decline an applicant, alter an obligation, or generate a notice a regulator can read. Those are a minority of an agent's actions and effectively all of its risk.

Where this points

The conclusion is not that the agencies erred. It is that the deferral changes what a well-run institution should do, and changes it in the opposite direction from the instinct. When a supervisor withdraws a specification while preserving every obligation attached to the underlying activity — safety and soundness, consumer protection and fair lending, third-party risk — the institution is not relieved. It is left to design the controls itself, with footnote 3 saying so in terms, and with footnote 1 confirming that supervisory action still follows unsafe or unsound practices stemming from insufficient management of model risk.

And the sequencing is the whole commercial point. The eventual specification, whenever it comes, will be written after a consultation, and consultations are answered by describing what firms already do. The institutions with a working authority record when that request for information opens will not merely pass an examination. They will be the material the specification is drafted against. That is a rare position to be in and it is available for a limited period, entirely on the merits of having built something.

What that something is, in structure, follows from the physical fact rather than from any regulation. If the authority decision is a function that is currently evaluated implicitly and discarded, the fix is to evaluate it explicitly and keep the result — to make the decision an object with the policy version pinned, the principal chain resolved, the context and evidence digested, the scope and constraints stated, and an expiry, produced at the instant of the decision rather than assembled afterwards by somebody reading logs. Every one of those fields already appeared, in prose, in the model inventory the 2026 revision just simplified away. The design question is what such a record has to contain to be worth anything to a reviewer who does not trust the system that produced it, how it is bound to the action so that it cannot be manufactured later, and what it costs in the hot path.

That is the subject of the companion piece to this one, Building the evidence layer the next examination will ask for, which is forthcoming on this site. This piece is the teardown, and a teardown that solves the problem is not a teardown. The claim here is narrower and I want to leave it standing alone: the specification you were waiting for has been withdrawn for your agentic systems, the obligations it attached to have not moved, the layer that would satisfy them falls between two exclusions and is governed by nothing, and the record you currently keep answers a different question from the one you will be asked.