What follows opens on a constructed illustration. It is not client work, not an incident report, and not a description of anything that occurred. It is assembled from primitives any delivery lead will recognise from a programme they have run, because the argument needs a trace concrete enough to check against your own accounts — and inventing an engagement to make a point would be exactly the behaviour this site argues against.
The programme runs eighteen months. Three squads, a shared orchestration layer, tool adapters into the client's system of record and its ticketing platform, a policy layer wired to the client's existing identity provider because re-platforming identity was explicitly out of scope. Agents go live in two functions — a procurement reconciliation flow and a claims triage flow — first in a shadow cohort, then in production behind an approval threshold. Every stage gate is passed. The deliverables schedule is discharged in full: architecture pack, control design document, test evidence, runbooks, four handover workshops, ninety days of hypercare. Acceptance is signed. The account moves to a smaller managed-service footprint and the build team is redeployed.
Nine months after that, the client's internal audit opens a routine review of the pilot period. The request that reaches the retained team is one sentence long, and it is not hostile. *Demonstrate what the deployed agents were permitted to do between the go-live date and the end of the pilot.* Not what they did — the audit team already has that. What they were permitted to do.
The retained team assembles what exists. The architecture pack describes the intended controls with unusual care: a tool allow-list per agent role, an approval threshold above a stated value, delegation limited to a single hop, scope reviewed at each gate. The policy repository has a clean commit history. The change-management system holds every approval. The runtime emits the ordered call sequence, the principal identity and session, timestamps to the millisecond, field-level before and after values, latency, retries and token spend. It is, by any reasonable standard, a well-run estate with excellent telemetry.
None of it answers the question. Not because a retention window closed — nothing was deleted. Not because sampling dropped it — this was not a sampled path. For any given action in that window, the authority under which it happened was never a record at any point in its existence. It was a value in a stack frame, computed correctly, consumed correctly, and released when the function returned. The design pack states what the controls were meant to be. The log states what occurred. Neither is a claim about what was in force at the instant of the write.
By the time the question arrives, the people who could have explained the difference are gone. The partner has rotated. The architect is eighteen months into another client. The two engineers who wrote the policy adapter have left the firm. The retained team is four people who joined at hypercare, and their honest answer — the intended design is documented and we have no artefact showing it was the design in force — reads to an audit committee as a control failure at the client. It is written up as a finding against the client. The remediation lands as an unbudgeted change request on the firm, at the worst possible moment, with none of the original staff available.
Nothing was missed against the scope. The scope was missed. Every line of the deliverables schedule was delivered and accepted. The evidence layer was not a line. It was not deprioritised, descoped or traded away in a commercial conversation — it never appeared in one, because at the point where a statement of work is negotiated nobody in the room is thinking in terms of an artefact that proves what a permission system decided.
This is a delivery-quality argument, not a compliance one, and the distinction matters for who should be reading. Firms that build inside somebody else's estate occupy an unusual position: they inherit whatever instrument governs the client, they carry almost none of it themselves, and the only obligation that is genuinely theirs is contractual — what the statement of work says will be handed over. That is a narrow surface. It is also the only surface on which this particular failure can be prevented, and it is almost never specified at the authority layer.
The objection, at full strength
There is a serious case that everything above is a category error, and it deserves stating properly before it is answered, because the version delivery leads actually hold is considerably better than the strawman.
*It was not scoped, so it was not owed. Acceptance is a legal event, not a courtesy. In the federal case the doctrine is stark: the Federal Acquisition Regulation's supplies clause provides that acceptance shall be conclusive, except for latent defects, fraud, gross mistakes amounting to fraud, or as otherwise provided in the contract*. A firm that delivers the schedule it signed, to the standard it signed, has discharged its obligation. Nine months of subsequent client curiosity does not retroactively expand a statement of work, and a profession that lets it do so has no commercial future.
*The firm cannot emit what the client's platform does not emit.* Platform selection was the client's, usually years before the engagement. The runtime, the identity provider, the policy engine and the logging pipeline were all inherited constraints. Asking a systems integrator to produce authority evidence from a stack that has no authority-evidence surface is asking it to fix the client's vendor choices at its own risk.
*The obligation is the client's and is not transferable anyway. The US banking agencies say so without hedging: a banking organization is responsible for conducting its activities in compliance with applicable laws and regulations, including those activities involving third parties. The use of third parties does not abrogate these responsibilities. India's data-protection statute is blunter still — the duty sits on the data fiduciary irrespective of any agreement to the contrary*. If the duty cannot be moved onto the firm, why is the firm designing to it?
*And the retrofit is revenue.* A remediation programme is billable work arriving in a warm account with an executive sponsor who now has a reason to fund it. On a pure margin reading, discovering the gap at month twenty-seven is better business than closing it at month three.
The strongest form of the objection is none of those, though. It is this: *the auditor's question is badly posed, and mature estates have always answered it by reference to design and configuration rather than per-action evidence.* Nobody asks a bank to prove, transaction by transaction, which entitlement matrix was loaded when a teller posted an entry. They ask for the access-control design, the current configuration, the change history, and a sample. That has been accepted practice for decades across systems far more consequential than an agent reconciling a supplier catalogue. Treating agents differently is an unfunded standard invented by people who write about agents.
Take all of that as granted. It is true, and the last one is nearly decisive. Here is precisely what each does not answer.
On scope: the contractual reading is right and commercially irrelevant. Acceptance ends the delivery obligation; it does not end the relationship, and the relationship is the asset the firm is actually managing. Worth noting too that the conclusive-acceptance sentence sits in the *supplies clause. The parallel services clause carries no such finality language; what it carries is a records provision requiring that inspection records be maintained and made available to the Government during contract performance and for as long afterwards as the contract requires* — a retention period with no floor at all, set entirely by the statement of work. A firm leaning on acceptance finality in a services engagement is leaning on a clause written for something else.
On platform constraint: true at the emission layer, false at the design layer. What a firm can always do — at no engineering cost — is *name what the estate cannot produce, in writing, at the point the estate is chosen*. A limitation recorded at acceptance is an asset. The identical limitation, undiscovered until month twenty-seven, is a finding. The difference between those two outcomes is a paragraph, and it is written before any code exists.
On non-delegability: it is exactly because the duty stays with the client that the firm is exposed. The client cannot discharge a duty it cannot evidence, so it discharges it through instructions, contracts and demands on whoever built the thing. The retained team is who that demand reaches. Non-delegability does not shield the builder; it guarantees the builder gets the call.
On retrofit-as-revenue: this is where the mechanism bites hardest, and it is not a margin argument. *The elapsed window cannot be backfilled at any price.* A remediation can instrument every future decision beautifully. It cannot manufacture evidence of what a policy engine decided in a quarter that has already closed, because the inputs to those decisions no longer exist. Whatever the firm sells into that finding, the finding stands for the period under review. The programme is judged on a window that is now permanently unevidenced, and that judgement attaches to the name on the architecture pack.
And on the last objection — the good one — the answer is that the proxy stopped holding, and it stopped holding because the client modernised. Design-plus-configuration works as a proxy for authority when configuration is slow-moving and enumerable. You can inspect an entitlement matrix because the matrix is a durable object that changes on a change ticket. That is not the model the client's own security architecture is being rebuilt around, and it is emphatically not the model an agent stack runs under.
The reference architecture for zero trust makes the point in its own tenets. Access is granted on a per-session basis, and it is determined by dynamic policy — including the observable state of client identity, application/service, and the requesting asset — and may include other behavioral and environmental attributes, with inputs the document lists as including time/date of request, previously observed behavior, and installed credentials. Read that as an engineer rather than as a policy document. It says the authorisation decision is a function over values that change between one request and the next by design. A decision function over inputs that mutate cannot be re-derived from an effects log written afterwards. The proxy did not fail because agents are special. It failed because the estate stopped being static, and the estate stopped being static because that was the recommendation.
What a log is, physically
Strip away tooling and a log is a plain object: an append-only sequence of records of effects. Three properties of that object matter here, and each is usually treated as too obvious to state, which is why the consequence is missed.
*The integrity guarantee is about the sequence, not the records.* Append-only means nobody rewrote entry 4,112 after the fact. It says nothing whatever about whether entry 4,112 contains the field somebody will need in nine months. A cryptographically perfect log of the wrong fields is a cryptographically perfect record of your own earlier assumptions. Integrity is orthogonal to sufficiency, and estates routinely buy the first while believing they have acquired the second.
*What gets written is what the emitter decided to write, under the emitter's model of what would later be asked.* That model is the schema, and a schema is therefore a forecast — a bet placed by its author about the shape of future questions. Most of the schemas an agent programme emits into were designed before agents existed, by people solving an operations problem, and they are excellent at the questions those people anticipated.
*And the load-bearing one: authority is not an effect.* An effect is a change in the world — a row updated, a payment released, a ticket closed. It has consequences and leaves residue. An authorisation verdict is a predicate over system state at an instant. Evaluating it changes nothing. It writes no row, sends no message, moves no byte outside the evaluating process. If nobody deliberately records it, it leaves no trace, because there was no trace to leave.
An allowed action and an unchecked action produce identical effects. The log cannot distinguish them, because the difference between them existed only inside the evaluator.
A denied action is starker still: it produces no effect, so a log of effects contains no evidence that a permission system exists at all. Two estates — one with a rigorous, well-tested policy layer, one with no authorisation whatsoever — emit effect logs that are indistinguishable in every field bearing on authority, for every request that succeeds. From the logs alone you cannot tell which one you built. That is not a criticism of anyone's implementation. It is what a log of effects is.
Five arguments, five different reasons the value does not survive
Authority resists being recorded as an afterthought because it is not a value. It is the output of a function with several time-varying arguments, and preserving the output without the arguments preserves almost nothing:
*authority = f(policy revision, principal and scope, request context, delegation depth, evidence at t) → verdict*
Each argument fails to survive the call for its own reason, and the reasons are different enough that no single fix addresses them.
Policy revision is not what the repository says. It is the bundle the evaluator had loaded when it decided. Under normal operation those agree. Under the conditions where the question actually gets asked — after a rollback, during a partial deploy, in the region that lagged, on the node that served a cached bundle through a failed fetch — they systematically disagree, and they disagree in the direction that matters, because the incident and the propagation failure tend to have a common cause. The repository is authoritative about what was committed. It is not authoritative about what ran.
Principal is not identity, and identity is the part that got solved. Identity is the part of this problem that has attracted real investment, and answering which agent instance acted is the question that investment has gone into. The auditor's question is about scope — what that well-identified principal was entitled to at that instant — and scope is mutable independently of identity. A role gains a permission in week six and loses it in week nineteen while the identity, the name and the audit trail all stay constant. Knowing who acted is progress on a different question.
Request context, in an agent stack, is frequently not an input at all. In a conventional service the attributes handed to the policy engine come from the request: a tenant identifier, a resource path, a source address — stable, reproducible, present in any decent trace. In an agent loop they are often outputs of an earlier stage: a value extracted from a retrieved document, a field the model summarised from a prior tool result, a risk score computed mid-run. The attributes that determined the verdict may not exist anywhere after the run ends, and may not be reproducible even by re-running the same task against the same corpus.
Delegation depth is standardised, carried, and explicitly excluded from the decision. This is the most interesting one, because a specification already anticipated it and then deliberately declined to use it.
RFC 8693, OAuth 2.0 Token Exchange, has defined the format since January 2020. Its *act claim provides a means within a JWT to express that delegation has occurred and identify the acting party to whom authority has been delegated, and a chain of delegation can be expressed by nesting one act claim within another, with the nested claims serving as a history trail. Then the same document constrains what a consumer may do with that trail: the consumer of a token MUST only consider the token's top-level claims and the party identified as the current actor*. The chain exists, it is standardised, it is carried in the token — and it is audit material by design, ruled out of the access-control decision on purpose. A Proposed Standard has been sitting on the delegation-history problem for over six years, and none of the specifications read for this piece wires it to a durable record of anything.
Now set that beside the specification the sector is currently building agent integrations to. The Model Context Protocol authorization specification, revision 2025-06-18, imposes strong requirements at the moment of use: servers MUST validate that access tokens were issued specifically for them as the intended audience, they MUST NOT accept or transit any other tokens, and a server MUST NOT pass through the token it received from the MCP client. Every one of those is correct — token passthrough is a real and well-understood vulnerability, and forbidding it is the right call. But the specification contains no requirement anywhere that a server emit, sign or retain any record of the authorisation decision it reached. Authority is checked and discarded, in the integration spec being adopted right now.
The correct security behaviour severs the chain the auditor wanted to walk. If the token does not cross the hop, the delegation relationship does not cross the hop either — not unless something at the boundary deliberately writes down that it happened. The upstream record names one principal, the downstream record names another, and nothing in either specification requires a correlation identifier joining them. For a firm whose delivery model routinely involves a subcontracted component or an offshore build pod, this is the hop where the trail ends.
Evidence at t is gone by construction. Where a verdict depended on a value — a credit exposure, a risk score, a retrieved document, a freshness check — the verdict is only meaningful alongside what that value was. Nine months later the exposure has moved, the score has been recomputed against a newer model, the document has been re-indexed. Reconstructing the input from current state produces a confident answer that is wrong, which is materially worse than producing none.
Why the discard is structural rather than a delivery lapse
A reader can accept all of the above and still hold that this is an oversight somebody will patch next quarter. I think that is wrong, and the reason is that at least six independent forces push toward the discard, none of which is a bug, and every one of which would have to be addressed deliberately.
*One: the evaluation is on the hot path.* Authorisation is called per tool call, and in an agent loop that is per step of something that may run hundreds of steps. Latency budgets are single-digit milliseconds. Emitting a durable, ordered, queryable record per evaluation adds a write to the highest-volume path in the system. Computing, returning and releasing is the correct engineering decision under those constraints, and no reviewer has ever flagged it as data loss.
*Two: the reference architecture prescribes the verdict as the record. This is the one that should stop a delivery architect mid-sentence. The zero trust architecture document describes its policy engine like this: the policy engine makes and logs the decision (as approved, or denied), and the policy administrator executes the decision*. A binary approved-or-denied entry is the standard's own prescribed record. The trust-algorithm inputs that produced it are not required to be retained anywhere. A team that implements the reference architecture faithfully, to the letter, ends up with exactly the gap this essay is about — and can point at the document to show it did the right thing.
*Three: the authorisation model underneath it returns a verdict and nothing durable. XACML 3.0 — the vocabulary the zero trust architecture borrows for policy decision and enforcement points — defines the policy decision point as the system entity that evaluates applicable policy and renders an authorization decision*, and specifies a response context carrying decision, status, obligations and advice. There is no normative requirement on the decision point to persist or emit the attributes it used. The request context is handed in, the decision is handed back, and the standard is indifferent to whether either survives the call.
*Four: the non-repudiation control in NIST SP 800-53 is an effects control. NIST SP 800-53 Rev. 5 — the US federal control catalogue a great many regulated estates inherit by reference — asks organisations, in its non-repudiation control, to provide irrefutable evidence that an individual (or process acting on behalf of an individual) has performed [organisation-defined actions]*. Read the scope carefully. It is about proving that an action occurred and who performed it. It is not about proving the action was authorised. When a client's control framework says non-repudiation is covered, what is covered is attribution of effects — which is precisely the half the estate already had.
*Five: the integration specification does not ask for it.* Covered above: strong audience-binding MUSTs at the moment of use, no requirement to record the decision, and an explicit prohibition on the passthrough that would have preserved the chain.
*Six, and this one is the sector's own: the evidence layer is nobody's deliverable. A services firm optimises against acceptance criteria, correctly and by design — that is what a professional delivery organisation is. The acquisition control in the same catalogue makes this explicit: nine categories must be included explicitly or by reference in the acquisition contract, among them security and privacy documentation requirements, allocation of responsibility or identification of parties responsible for information security, privacy, and supply chain risk management, and acceptance criteria*. Those three lines are where an authority-evidence deliverable either exists or does not. If it is not in them, no delivery organisation on earth will produce it, and none should be criticised for that. The failure is upstream of delivery.
Six forces, none an error, all pointing the same way. That is what structural means. A gap sustained by one oversight gets patched when one engineer notices. A gap sustained by a latency budget, a reference architecture, an authorisation model, a control catalogue, an integration specification and a contracting practice gets patched only when somebody decides it is a deliverable.
Four reconstructions, and where each one breaks
When the question arrives, teams try to answer it. Four strategies are available, and each fails on a different mechanism. Working through them is what converts we should have logged more into an understanding of why more logging of the same kind would not have helped.
- *Reconstruct from the policy repository.* Read the commit history, find the revision in force on the date, replay it. Fails on propagation: what the repository held and what the evaluator loaded are different facts, and they diverge exactly under the conditions that produce the incidents worth auditing. The repository is a record of intent, which is what the design pack already was.
- *Reconstruct from the principal's current permissions.* Look up what the agent identity is entitled to today and assert it held that scope then. Fails on drift, and fails in the most dangerous available way: it returns a confident, well-formatted, plausible answer that is wrong whenever anything changed in nine months. The audit team accepts it because it has the shape of evidence. This is the failure mode that should worry a delivery principal most, because it is the one that does not look like a failure and is the one most likely to be produced under time pressure by people who joined at hypercare.
- *Reconstruct from the delegation chain.* Walk back from the acting principal to the human who granted the authority. Fails at the first hop where a token was exchanged rather than forwarded — which, under the integration specification's own rules, is every hop, by mandate. The chain format exists and is defined as informational rather than decisional; the correlation identifier that would let two records be joined across the exchange is required by neither specification.
- *Reconstruct from the effect diff. Argue that since we can see exactly what changed, the authority is implied. Fails on arithmetic. Effects underdetermine authority: many distinct policy states permit the identical write. Recovering some authority existed from the write succeeded* is trivially true and answers nothing that was asked, because the question distinguishes a narrow reviewed grant from a blanket one, and both produce the same bytes.
The general form is a projection. The effect record is the authority tuple projected onto its observable coordinates. Projections lose information and are not invertible. No improvement in retention, integrity, completeness or query performance changes that — a discarded dimension cannot be recovered by storing the surviving ones more carefully. This part of the argument is not an engineering opinion.
Three things the retained team can actually run
Nothing here is a proposal, and none of it is the architecture — this is the teardown. The first tab shows that many policy states produce one effect record, which is the whole non-invertibility claim reduced to something you can execute. The second walks a delegation chain across a subcontracted hop and shows where it terminates. The third models a deliverables schedule against the evidence classes an audit will ask for, which is the check worth running on a live statement of work this week. Paste any of them into a TypeScript playground.
A deliberately small model. Four policy states are all consistent with the single effect the log recorded, and they differ in exactly the way an audit question cares about: a narrow reviewed grant and a blanket one are indistinguishable downstream. The function does not fail to find the answer — it correctly finds that more than one answer exists, which is a different and much worse result.
type PolicyState = {
readonly revision: string;
readonly scopes: readonly string[];
readonly approvalThreshold: number | null; // null means "no threshold configured"
readonly maxDelegationHops: number;
};
type EffectRecord = {
readonly action: string;
readonly resource: string;
readonly amount: number;
readonly actorId: string;
readonly ts: string;
};
/** The four states the estate could plausibly have been in during the pilot window. */
const CANDIDATES: readonly PolicyState[] = [
{ revision: "r-118", scopes: ["catalogue:write"], approvalThreshold: 5_000, maxDelegationHops: 1 },
{ revision: "r-119", scopes: ["catalogue:write"], approvalThreshold: 50_000, maxDelegationHops: 1 },
{ revision: "r-120", scopes: ["catalogue:write", "ledger:write"], approvalThreshold: null, maxDelegationHops: 3 },
{ revision: "r-121", scopes: ["catalogue:*"], approvalThreshold: null, maxDelegationHops: 9 },
];
/** Would this state have permitted the effect the log recorded? */
function permits(state: PolicyState, effect: EffectRecord): boolean {
const scoped =
state.scopes.includes("catalogue:write") || state.scopes.includes("catalogue:*");
const underThreshold =
state.approvalThreshold === null || effect.amount <= state.approvalThreshold;
return scoped && underThreshold;
}
const observed: EffectRecord = {
action: "record.update",
resource: "catalogue/supplier-price",
amount: 1_200,
actorId: "agent-recon-07",
ts: "2025-03-04T09:41:22.118Z",
};
const consistent = CANDIDATES.filter((state) => permits(state, observed));
console.log(consistent.map((s) => s.revision)); // ["r-118","r-119","r-120","r-121"]
console.log("policy states consistent with the log:", consistent.length); // 4
// The reviewed grant (r-118: a 5,000 threshold, one delegation hop) and the blanket one
// (r-121: wildcard scope, no threshold, nine hops) are indistinguishable from this record.
// That distinction is the entire content of the audit question.The three files are independent and none of them depends on the others. The schedule in the third tab is illustrative — replace it with the one in front of you, which is the only version of the exercise that is worth anything.
The contract surface, clause by clause
A firm that builds inside somebody else's estate carries almost no direct regulatory obligation of its own. It inherits the client's, and it converts that inheritance into a set of contractual promises. So the useful exercise is to read the provisions that actually govern the handover and ask, of each, what class of record it names. *Say which instrument, and say whose it is* — the instruments below are the client's, and the firm meets them only through the words in the statement of work.
For a US bank client, the interagency guidance issued by the Federal Reserve, the FDIC and the OCC in June 2023 is the frame everything else hangs from. It sets out four provisions that matter here. On retention, it names a whole contract section — responsibilities for providing, receiving, and retaining information — covering the banking organisation's ability to access its data in an appropriate and timely manner, access to the third-party's data and any supporting documentation, and how data may be shared with regulators in a timely manner as part of the supervisory process. Every enumerated item is data, documentation or a report. Nothing in the section names an authorisation decision as a class of retained information.
On inspection, it extends the audit right through the prime to whoever the prime brought with it: generally, a contract includes provisions for periodic, independent audits of the third party and its relevant subcontractors, consistent with the risk and complexity of the third-party relationship, and it asks whether the contract provisions describe the types and frequency of audit reports the banking organization is entitled to receive. That is the clause under which somebody eventually asks a services firm to produce evidence it never designed anything to emit.
On subcontracting, it names the structural problem for routers precisely. Subcontracting can result in risk due to the absence of a direct relationship between the banking organization and the subcontractor, further lessening the banking organization's direct control of activities, and it recommends contractual obligations such as reporting on the subcontractor's conformance with performance measures, periodic audit results, and compliance with laws and regulations, plus provisions stating the third party's liability for activities or actions by its subcontractors. If a build pod, an offshore centre or a specialist partner sits inside the delivery model, the firm has already accepted a liability whose evidence base is whatever that partner happens to emit.
And on supervision, it recommends the contract stipulate that performance is subject to regulatory examination and oversight, including appropriate retention of, and access to, all relevant documentation and other materials — footnoted to the statutory examination authority at 12 U.S.C. 1464(d)(7)(D) and 1867(c)(1). Note where the weight falls. The retention duty a services firm actually carries is scoped by the phrase all relevant documentation, which is defined nowhere and is settled, in practice, in the statement of work.
Two more instruments sharpen the picture, both from outside financial services and both showing what a regulator writes when it genuinely wants evidence preserved. The defense supplement clause on safeguarding and incident reporting requires a contractor to preserve and protect images of all known affected information systems … and all relevant monitoring/packet capture data for at least 90 days, and requires the clause be included in subcontracts without alteration, except to identify the parties — the flow-down mechanism by which a firm's own subcontractors inherit the client's instrument verbatim. It is a real, specific, enforceable preservation duty. It is scoped entirely to effects: images and captured traffic. Even where the drafting is at its most demanding, the object preserved is what happened.
The federal acceptance clauses complete it. The supplies clause makes acceptance conclusive except for latent defects, fraud and gross mistakes amounting to fraud. The services clause carries no equivalent finality and instead requires inspection records be maintained and made available … for as long afterwards as the contract requires. Read together, the asymmetry is instructive rather than comforting: in services, the retention obligation has no floor, and the words that set it are the ones the firm drafted.
For an Indian client the analysis is different and, for a firm headquartered in India, closer to home. The Digital Personal Data Protection Act 2023 places the duty on the data fiduciary and makes it non-contractible: a fiduciary shall, irrespective of any agreement to the contrary or failure of a Data Principal to carry out the duties provided under this Act, be responsible for complying with the provisions of this Act in respect of processing undertaken by it or on its behalf by a Data Processor. It then says the fiduciary may involve a processor only under a valid contract. An IT services firm is typically the processor. Its entire obligation is that contract, and the statute has already ruled that the duty cannot be pushed back onto it.
Two further sub-sections turn that into build requirements. The fiduciary shall protect personal data in its possession or under its control, including in respect of any processing undertaken by it or on its behalf by a Data Processor, by taking reasonable security safeguards to prevent personal data breach; and it must cause its Data Processor to erase any personal data that was made available by the Data Fiduciary for processing. Both are duties the client can only discharge through instructions the firm's build has to be capable of both executing and evidencing. An erasure instruction that an agent estate executed, with no artefact showing under what authority and against which scope, is a duty performed and unprovable.
For an Indian bank or NBFC client, the RBI's 2023 directions on outsourcing of information technology services come closest of anything read for this piece to naming the thing. The outsourcing agreement must provide effective access by the regulated entity to all data, books, records, information, logs, alerts and business premises relevant to the outsourced activity; there is a right to audit the service provider including its sub-contractors; sub-contracting requires prior consent; the service provider is contractually liable for the performance and risk-management practices of its sub-contractors; and outsourcing shall not diminish the regulated entity's own obligations. That is the most evidence-forward drafting in the set. It still says logs.
For a Gulf client the discipline is the same and the citation is not mine to give: the instrument is whichever supervisory framework and sovereign-programme commitment the client is actually bound by, and a firm that cannot name it in the statement of work has not finished reading the engagement. What does not change across any of these jurisdictions is the shape of the exposure. Every instrument puts the duty on the client, gives the client rights against the firm, and leaves the definition of the evidence to a document the firm writes.
One more piece of context matters for a bank client specifically, because it removes the last place a delivery team might have expected the specification to come from. The revised interagency model-risk guidance of April 2026 — OCC Bulletin 2026-13, Model Risk Management: Revised Guidance, with the parallel Federal Reserve issuance SR 26-2 — supersedes SR 11-7 and SR 21-8. In the guidance's own text — footnote 3 of the interagency guidance attachment issued jointly by the Board of Governors of the Federal Reserve System, the FDIC and the OCC, which the OCC publishes as Bulletin 2026-13 and the Federal Reserve as SR 26-2 — generative and agentic AI models are not within the scope of this guidance; the bulletin also states that it does not set forth enforceable standards or prescriptive requirements. Read it as a deferral rather than an exemption. Every obligation attached to the underlying activity — safety and soundness, consumer protection, third-party risk — is untouched. What has been withdrawn is the framework that would have specified the controls. So for the moment nobody is going to hand a delivery team an agent-control specification. The specification is the contract, and the contract is the deliverables schedule.
What the field already knows, and where it stops
It would be dishonest to present this as unexplored ground. Most of the intellectual work has been done, in the open, by standards bodies whose documents are freely available. What has not happened is anyone requiring the artefact.
The architectural separation the argument wants already exists in the reference model. The policy decision point is broken down into two logical components: the policy engine and policy administrator, and those sit on a control plane while the enforcement point that performs the action sits on the data plane. The split between deciding and doing is not something an essay needs to propose. It is the standard's own model. What is missing is a durable object that crosses between the planes — the decision passes across as an instruction to act, and nothing is required to persist carrying the reasoning with it.
The materials for such an object are also standardised and mostly unused. RFC 8693's *act and may_act* claims give a signed, nestable, verifiable representation of delegation with an explicit history trail, in a Proposed Standard from January 2020. XACML's response context has room for obligations that a decision point can attach to a verdict, though they are discharged by the enforcement point rather than persisted as evidence. Between them, a signed statement of who delegated what to whom, under which policy, is not a research problem. It is an assembly problem, and the assembly is the companion piece's subject rather than this one's.
Then the honest negatives, which matter more than the positives for a reader deciding whether to act on this.
*No standard requires an authorisation receipt.* Across everything read for this piece — the zero trust architecture, the control catalogue's audit and acquisition families, XACML 3.0, RFC 8693 and the Model Context Protocol authorization specification at revision 2025-06-18 — not one imposes a requirement to emit or retain a signed artefact of the authority decision. The nearest constructs are obligations discharged by the enforcement point, a non-repudiation control scoped to proving an action occurred, and delegation claims carried in a token but explicitly excluded from the access-control decision. The absence can be asserted. A standard that requires the thing cannot be cited, because there is not one.
*And the headline number does not exist.* The obvious question — what fraction of an agent's actions can have their authorising decision correctly reconstructed from logs alone, as a function of elapsed time and policy churn — has no primary publication behind it that I could locate. No standards body, regulator or vendor publishes a reconstruction-rate measurement, and no benchmark defines the metric. That is itself a finding, and a more interesting one than a statistic would have been: nobody quotes a number because nobody has instrumented the experiment. Any figure encountered in the wild for this should be treated as fabricated until somebody produces the primary source. I have not run the experiment either, and this piece would be improved by someone running it.
The limits of this argument
Several things could be true that would weaken or overturn what is above, and a teardown that does not name them is advocacy.
- *The trace is constructed.* The opening is an illustration built from recognisable primitives, not a case, and I am making no claim about how often it happens. The argument is about mechanism: if the estate never materialised authority, the question cannot be answered from it. How frequently that question gets asked is an empirical matter I have no data on.
- *The cost claim is a mechanism claim, not a measured one.* I can explain why retrofitting is more expensive than designing in — no original staff, no test harness, production change control, and an elapsed window that cannot be backfilled at any price. I cannot give you a ratio, and inventing one would poison everything else here.
- *Emission could close without any standards change.* If a mainstream policy engine or agent runtime shipped a signed decision record carrying its inputs, default-on, the structural argument weakens considerably in one release. Nothing in any specification prevents a vendor doing this tomorrow; the obligations discussed above are permissive, not prohibitive. I would regard that as the best available outcome and it would date this piece usefully.
- *The audit practice might simply not converge on the question.* If review teams continue to accept design-plus-configuration as sufficient for agent estates, as they largely have for conventional ones, the exposure stays theoretical. My answer is the dynamic-policy argument above, and it is an argument rather than an observation. Someone with a large sample of actual agent-estate audit findings could falsify it, and I would want to see that sample.
- *The instruments read here are US federal, US banking and Indian.* That is where the argument is anchored and it is not a survey of every regime. A firm working under a supervisory framework not read here should assume the analysis needs redoing against the actual text rather than assuming it transfers.
- *The clause-level readings are as good as my sources.* Section and paragraph references are given so they can be checked, and where I could not confirm a specific numbering from raw text — the RBI directions in particular — I have quoted the substance and named the direction rather than printing a clause number I could not verify.
Where this goes
The shape of the answer is not mysterious, and stating the shape is as far as this piece goes deliberately. Something has to be emitted at the moment of decision, carrying the arguments to the authority function rather than only its output, bound to the effect it authorised so the two cannot drift apart, signed so it means something to a party that trusts neither the firm nor the runtime, and retained on the client's own retention terms rather than the runtime vendor's. That is a receipt. Every component of it is a settled standard in some adjacent domain.
The hard part is not the design. The hard part is that a services firm has to do it inside an estate it does not own, cannot re-platform, and will hand back — with an identity provider chosen years ago, a runtime chosen by a different part of the client, and a change-control regime that treats the policy layer as production. That is a genuinely difficult engineering problem with several unattractive answers, and it is the subject of the companion piece on this site, Building a permission receipt layer inside an estate you do not own.
What can be done before that piece exists costs nothing and takes an afternoon. Open the statement of work for whatever is currently in build. Find the acceptance criteria. Read the deliverables schedule and ask, of each audit question a reviewer might pose about the pilot period, which named deliverable answers it. Where nothing does, write the sentence — either specifying the artefact or recording, in writing and at signature, that the estate as selected cannot produce it. Both outcomes are defensible. Only one of them is available nine months after handover, and it is neither.