Start with the chain, because the argument is not abstract and the abstraction is what usually gets it dismissed.
A first notice of loss arrives — water damage, a burst supply line, a homeowner with photographs. An intake agent classifies the loss and opens the claim. A coverage agent reads the policy in force on the date of loss, the endorsements, the deductible. An estimation agent prices the repair from the photographs and a contractor database. A settlement agent applies depreciation, subtracts the deductible, posts a payment to the policyholder for less than the estimate, and drafts the letter that explains why. Five hops. One of them moves money to a policyholder. One of them produces an adverse communication — the letter that says the recoverable amount is smaller than the loss.
The telemetry is good. Spans, parents, latency per hop, tool arguments and results attached. If the estimation agent had timed out, an engineer would find it in ninety seconds.
Then, in a market conduct examination — or in the file review that precedes one — someone asks a question that sounds administrative and is not:
Which adjuster decided to settle this claim below the estimate, and was that within their authority?
The record answers a neighbouring question with total confidence. The payment was posted by a service principal — call it svc-claimsplatform-prod. That principal holds an entitlement to the payments endpoint. The entitlement has an owner of record, was provisioned through the joiner-mover-leaver process, and was recertified nine weeks ago by a named person who is a director in claims operations. Every link in that sentence is real, evidenced and testable.
None of it answers the question that was asked, and in insurance the question has a second clause that makes the failure worse than banking's version of it. The examiner is not only asking for a person. They are asking for a person within an authority — because settlement authority is how this sector has expressed per-decision permission for as long as claims have been paid. A named adjuster, a monetary limit, a recorded referral above the limit. The chain's record contains no name, no limit, and no referral, because every hop ran with the same authority: all of it.
What actually closes the gap, in practice, is a paragraph. Somewhere in the AIS Program — the written program the NAIC Model Bulletin asks insurers to maintain — there is a sentence of the form: the claims automation platform operates under the delegated authority of the Vice President, Claims, who has approved its entitlement set and retains accountability for actions taken under it. That sentence is written by a person, about a system, in a document the system has never read and cannot contradict. It may even be true. It is an assertion, not evidence — and the distance between those two is the entire discipline a market conduct examination exists to enforce everywhere else in the file.
The five-hop claims chain above, the service principal name, the AIS Program paragraph and the recertification detail are a constructed illustration, assembled from patterns documented in public specifications, published regulatory guidance and vendor documentation. It is not a report of any real incident, carrier, policyholder or deployment, and no part of this piece describes client work.
The objection, stated properly
There are two strong responses to what I have just written, and they come from different people. The first comes from inside the carrier, and it is the better one. The second comes from the regulatory reading, and it is the more comfortable one. Both deserve full strength.
The response from inside the carrier goes like this. Claims is one of the most accountability-dense functions in financial services, and it did not get that way decoratively. Settlement authority is graduated and personal, and the schedule is maintained, reviewed and enforced in the claims platform itself. Referrals above authority are recorded events with named approvers. The claim file — the sector's native evidence object — carries every note, every reserve change, every payment, timestamped and attributed, precisely because market conduct examiners read claim files and unfair claims settlement practices law attaches to what they find. Large-loss committees exist. Reinsurers audit the files too. And since the Model Bulletin, there is a written AIS Program with a named senior owner. In a sector where an examiner can pull five hundred files and count the defects, nobody forgot to write down who decides.
So the natural reading of my complaint is that I have described a carrier without a claims organisation, which is not a carrier. Add the field. Stamp the accountable executive's identifier on the span, propagate it, and the record now names a person at every hop.
The second response is the regulatory one. Nothing binding asks the question yet. The Model Bulletin is guidance: it interprets law that already exists — unfair trade practices acts, unfair claims settlement practices acts, market conduct examination authority — and creates no new legal obligation of its own. The Evaluation Tool is a pilot instrument; it has not been adopted, and pilots change. Adoption of the bulletin itself stands at more than twenty jurisdictions, not fifty-plus, so a national carrier faces a patchwork rather than a mandate. And the loudest thing any US financial regulator has said about agentic systems — footnote 3 of the April 2026 interagency model risk guidance — placed them expressly out of scope. On the most economical reading: build the thing, monitor it sensibly, and revisit when someone writes a rule.
I want to concede both of those at full strength before answering either, because each is right about something, and the thing each is right about is not the thing at issue.
Settlement authority is real — and it is the strongest version of the objection, because it proves the sector already built the missing object for humans. An authority limit is two things at once: a standing bound (this person may settle up to this amount) and a per-decision event generator (anything above the bound produces a referral — a request, an approver, a timestamp, a decision). The claims organisation is constructed out of exactly the request-approve-record triples that make a control auditable. Now look at what the chain replaced it with: a service principal whose entitlement is the union of everything any hop might ever need. There is no bound, so nothing is ever above the bound, so the event generator never fires. The chain did not merely fail to record its referrals. It is architected so that a referral can never occur.
Stamping the executive's identifier on the span writes down a fact that is not true at that hop. The VP of Claims did not decide to settle this claim below the estimate. They approved an entitlement set for a platform, in a change record, some months previously. Propagating their identifier onto the fourth hop of a chain they have never seen asserts a decision that did not occur — and it is worse than the paragraph in the AIS Program, because it now looks like data. An examiner who finds an identifier in a field will reasonably assume it was recorded because something happened. Manufacturing that assumption is not a control; in a claim file it is closer to the thing file reviews exist to catch.
Guidance, pilot and patchwork do not add up to absence — because the bulletin interprets law that already binds, and the pilot is that law's examination being operationalised. The Model Bulletin's legal theory is not that new obligations exist. It is that decisions affecting consumers which are made or supported by AI systems must comply with the same unfair trade practice and claims settlement standards that have always applied, and that insurers should expect regulators to ask for the governance evidence — the written program, the inventory, the third-party oversight — in examinations and investigations. The Evaluation Tool is what 'expect regulators to ask' becomes when it is turned into a document: a structured instrument for examiners, piloted in twelve states, iterated to version 4.0 mid-pilot. And the patchwork cuts the other way from the way it is usually argued: a carrier writing in forty states does not get to build to the most permissive one. It builds to the strictest examination it can be subjected to, which means variance raises the effective bar rather than lowering it.
Which is the argument, compressed. In banking, the agencies withdrew the specification and left the institution to determine its own controls. In insurance, the regulators kept the law exactly where it was and are building the examination instrument first. Those are opposite postures with the same consequence for an agent chain: the record it produces will be asked for, and the record cannot answer.
What the instruments actually are, and when the worksheet arrives
It is worth being precise about the stack, because it is small, dated, and mis-cited in ways an insurance reader will catch.
First, the bulletin. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurance Companies was adopted by the NAIC membership on 4 December 2023, and by mid-2026 has been adopted in more than twenty jurisdictions. Even its title requires care: much of the adoption-period material circulates the rendering 'by Insurers,' while the NAIC's own page renders 'by Insurance Companies' — this piece prefers the NAIC rendering. A model bulletin is not a model law: it creates no new obligation, and adopting states issue it as guidance interpreting law already on their books. What it asks for is concrete: a written program for the responsible use of AI systems, governance with defined roles, risk management and internal controls proportionate to the decision's consumer impact, and oversight of third-party AI systems and data — including the expectation that the insurer can produce this on request. The operative consequence sits in that last clause. The bulletin tells insurers, in advance, that the AIS Program and its documentation are examinable material.
Second, the Tool — the part of the stack that moves the question from 'could be asked' to 'is being asked, from a script.' The NAIC AI Systems Evaluation Tool is the examination instrument being built to operationalise the bulletin: structured material state examiners use to evaluate an insurer's AI governance, developed under the NAIC's Big Data and Artificial Intelligence Working Group. A twelve-state examiner pilot runs January through September 2026. The instrument reached version 4.0, discussed in a public session on 1 June 2026 — an iteration count that tells you the questions are being sharpened against real examinations, mid-pilot. Consideration for adoption is expected at the Fall 2026 National Meeting. Expected is the right verb and I will not improve on it: the pilot could conclude in revision, delay, or a weaker instrument, and the falsifiability section below takes that seriously.
Two honesty notes before building anything on this. I am describing the Tool's function — structured worksheets in examiners' hands — from the NAIC's public committee material about the pilot, not quoting its internal question text, which is pilot-stage work product. And the sentence that matters does not depend on the wording of any question: a published examination script, whatever its final text, cannot be answered with a bridging paragraph, because a scored worksheet has fields, and a field wants a name, a date, a limit, a record. Prose is what fills the margins of an exam script, not the blanks.
Third, New York. Insurance Circular Letter No. 7 (2024), issued by the Department of Financial Services on 11 July 2024, addresses the use of artificial intelligence systems and external consumer data and information sources — ECDIS, in the letter's vocabulary — in insurance underwriting and pricing. Its structure is the part that transfers to the authority question. The letter does not ask the insurer to assert that its data and models are fair. It asks the insurer to be able to demonstrate it: comprehensive assessment, quantitative testing appropriate to the use, board and senior management oversight of a governance framework, documentation an examiner can follow — and, pointedly, it does not allow reliance on a vendor to discharge any of this. An insurer using a third party's data or model remains responsible for demonstrating what that data and model did. Demonstration versus assertion is the entire distance between an authority record and a bridging paragraph, stated by a regulator two years before the agentic version of the question arrived.
The lineage consequence deserves its own sentence, because agent chains make it quantitatively worse. CL 7's demand is per-use: for this underwriting or pricing outcome, what external data entered it, from which source, and what testing supports its use. A retrieval-augmented agent chain multiplies external data touches — every retrieval hop is a potential ECDIS use — while the ambient-authority pattern erases the record of which source fed which decision. The same architecture that cannot name the adjuster also cannot name the data. It is one defect appearing on two axes: the chain does not carry provenance for authority, and it does not carry provenance for inputs, for the same structural reason.
Fourth, India — where the same argument arrives in drafting-gap form. IRDAI constituted a seven-member working group on artificial intelligence on 18 June 2026, with a three-month window to report. That window closes as the NAIC's examiner pilot ends, which puts the two markets in an instructive alignment: one regulator is finishing the pilot of an examination instrument, the other is starting to draft its framework. For an Indian insurer — or the Indian arm of a global carrier — the consequence is the same one this practice argues in banking: what the industry has demonstrably built by the time the working group reports is the material the framework will be drafted against. A claims chain whose authority record terminates in a service account is a poor exhibit. A chain whose every settlement carries a named, bounded grant is the exhibit a working group cites.
Fifth, the Gulf, in one honest sentence: as of 18 August 2026 I can verify no insurance-specific AI instrument in any GCC jurisdiction — the region's 2026 moves are structural and cross-sectoral rather than supervisory guidance for insurers — and anyone who cites one to you should be asked for the document.
Last, the banking context, because insurance readers inside financial groups will meet it. On 17 April 2026 the banking agencies issued revised interagency model risk management guidance — Fed SR 26-2 / OCC Bulletin 2026-13 — and footnote 3 of the shared interagency document places generative and agentic AI models outside its scope while returning responsibility for their governance to the institution; the consultation the agencies promised had, as of 18 August 2026, not been published. None of that governs an insurance claims chain, and this piece does not pretend it does. It matters here for one reason: a bancassurance group or an insurer with a banking affiliate now faces both postures at once — a deferral on one side of the house and an examination pilot on the other — and they land on the same architecture. Building the authority record satisfies both; the deferral does not excuse the worksheet.
The claim file was the answer, and the chain routes around it
Everything above is context. This section is the argument, and it is a claim about what exists in memory rather than about what anyone failed to log.
Insurance already owns the evidence object every other sector is missing. The claim file is a per-decision record constructed to be examined: every note attributed, every reserve change dated, every payment tied to an authority, maintained in that shape for decades precisely because examiners, courts and reinsurers read it adversarially. The agent chain does not corrupt the claim file. It routes around it. The chain's decisions happen in orchestration — in model calls, tool dispatches and handoffs the file never sees — and what lands in the file is the output: a payment amount, a letter, a closing code. The file records that the decision happened. The record of how, and on whose authority, was never created anywhere, because in the prevailing pattern authority moves between agents as ambient context. Three mechanisms account for nearly all of it, and the property they share is that none of them is an event.
- A shared session. The chain runs inside one session object authenticated once at the boundary — the platform's, not any adjuster's. Each agent reads from it. Nothing is passed, because the session is simply in scope; there is no moment at which the estimation agent gives the settlement agent anything, so there is no moment at which a grant could have been recorded.
- An inherited credential. A token or a pre-constructed client sits in the process environment, built at startup. The settlement agent does not obtain it. It has possessed it since the process began, on identical terms to every other agent — which is why its payment is indistinguishable, in the record, from any other hop's read.
- Context carried forward. Instructions, tool definitions and accumulated conversation state cross the handoff. What changes is what the receiving agent knows. What never changes is what it may do, because capability was not modelled as a thing that travels — or narrows.
Compare that with how the claims organisation produces its evidence. A referral above authority is an event: a request, a named approver, a timestamp, a decision. A reserve change is an event. A large-loss committee sign-off is an event. The claim file is auditable because the process is constructed out of request-approve-record triples, and a triple leaves a trace by construction. The agent handoff produces no request and no approval, and therefore nothing to record. It is not that the record is thin. The shape of thing the file knows how to hold was never instantiated.
This is why it is not a missing field. A missing field implies there was a value nobody wrote down. Here there is no value. When the settlement agent posts a payment under a credential minted before the claim existed, the only principal identifier that could truthfully be attached to that span is the one already attached to every other span — which distinguishes nothing. Writing it five times produces one fact copied five times, and the copies say nothing about what happened in between.
And it is why better instrumentation does not reach it. Turn sampling off. Retain everything at full fidelity. The chain is now perfectly recorded and the examiner's question is exactly as unanswerable as before, because the fidelity is fidelity to a call graph. A call graph is recoverable because each edge is an invocation, and an invocation is an event the code can observe while it happens. An authority graph is not recoverable because each edge would be a grant, and in this pattern no grant is ever created. The general form of that argument, made against the published specifications rather than against an examination regime, is in a separate piece for platform engineers; what matters for a carrier is the narrow version: the object your claim file needs does not exist, and a file built to evidence decisions cannot evidence a decision that was never made.
What fails, at mechanism level, in a carrier
Generic arguments about agent risk are cheap. Here is what specifically breaks, in the terms this sector uses, and the instrument each failure runs into.
The authority ladder, and the referral trail that vanishes with it. The settlement-authority schedule is the claims organisation's core financial control, and it has a property the entitlement model lacks: it is denominated in the decision's own units. Dollars, per claim, per person. The chain's entitlement is denominated in endpoints — the platform may call the payments API — which is a statement about capability, not about any decision. When an examiner samples paid files and asks the routine question, within whose authority was this payment, the human files answer in one lookup. The chain's files answer: within the platform's, which is to say within everyone's and no one's. Note that this failure does not depend on the Evaluation Tool being adopted. It is a defect against the carrier's own existing control framework, visible in any file review that thinks to ask.
ECDIS lineage, per decision. Circular Letter No. 7's demand structure — demonstrate, per use, with testing, without hiding behind the vendor — presupposes the insurer can say which external data entered which decision. A retrieval hop that pulls a third-party score, a property-characteristics record or a contractor-pricing feed into the context window is an external-data use whose linkage to the eventual settlement exists only in the conversation state, which is precisely the thing the record does not preserve as an object. The carrier is left demonstrating at the level of the system — we use these sources, we tested them annually — when the letter's grammar is the level of the decision.
Third-party hops the carrier cannot see inside. The Model Bulletin's third-party section expects the insurer to oversee AI systems and data acquired from vendors, and CL 7 refuses to let vendor reliance discharge the obligation. Transfer that to a chain in which one hop is a hosted vendor model or an external claims-automation service. The carrier cannot see inside the vendor's hop — that is what proprietary means. Under both instruments, opacity is not a reason the obligation lapses; it is a reason the carrier needs a record of what authority and what data crossed the boundary. That record is exactly the object the handoff never creates, and no vendor questionnaire retroactively creates it.
Two operational consequences follow, and they are the ones that hurt when something goes wrong.
The first is revocation. An entitlement can be pulled from a service principal in seconds. What cannot be done is revoking authority from a chain, because there is no chain object to revoke from — only the principal, and revoking the principal stops everything the platform does, including the claims it was handling correctly, in the middle of a catastrophe surge if that is when the defect surfaces. The carrier chooses between a blunt outage and an unbounded exposure, and both are answers a chief claims officer will find unacceptable.
The second is population scoping, and in insurance it has a name: re-adjudication. When a settlement logic defect is found, the first question is which claims were touched by the defective authority, and the answer has to survive being audited by a regulator who can order restitution with interest. If authority is ambient, the defensible population is every claim the service principal touched in the window — dramatically larger than the true population. The carrier either re-adjudicates at that scale and absorbs the cost, or narrows the population by judgement it cannot evidence and hopes the examination agrees. That is a decision made under uncertainty the architecture created and did not have to.
The examiner's question, asked against the record that exists
Two files, in the order a file reviewer would meet them. The first transcribes the shape of the record an agent claims chain actually produces, alongside the settlement-authority schedule the carrier already maintains for people. The second asks the examination question against it, and the type system makes the honest answer unavoidable: the function that should return a named adjuster within an authority has no way to construct one, because no input carries either the name or the limit.
Written from the shape of a conventional agent trace plus the two things a carrier can genuinely produce alongside it: the entitlement snapshot for the service principal, and the human settlement-authority schedule. Note where the authority schedule lives: entirely outside the chain record. Nothing in the run references it, because no hop ever consulted it.
/** A service identity. Stable, long-lived, provisioned through joiner-mover-leaver. */
export interface ServicePrincipal {
readonly kind: "service";
readonly id: string; // e.g. "svc-claimsplatform-prod"
/** Owner of record in the CMDB. An owner is not the adjuster of any claim. */
readonly ownerOfRecord: string;
readonly lastRecertifiedIso: string;
}
/** One rung of the carrier's settlement-authority schedule — maintained for PEOPLE. */
export interface AuthorityGrantHolder {
readonly directoryId: string;
readonly displayName: string;
readonly role: "adjuster" | "senior_adjuster" | "claims_manager" | "head_of_claims";
/** Monetary limit, in cents, per claim. The unit of the decision itself. */
readonly settlementAuthorityLimit: number;
}
/** One hop. Everything here observes an invocation that happened. */
export interface Hop {
readonly spanId: string;
readonly parentSpanId: string | null;
readonly operation: "invoke_agent" | "execute_tool";
readonly component: string; // "intake" | "coverage" | "estimation" | "settlement"
readonly startedIso: string;
/** Present only on the hop that reached a system of record. */
readonly effect?: {
readonly endpoint: string; // "POST /claims/:id/payments"
readonly amountCents?: number;
};
}
/**
* The run. The principal sits on the run, not the hop, because it is a property of the
* process: every hop executes under the identical value. The authority schedule is not
* referenced anywhere in this type — which is the finding, expressed as a type.
*/
export interface ClaimsChainRecord {
readonly claimId: string;
readonly correlationId: string;
readonly principal: ServicePrincipal;
readonly hops: readonly Hop[];
readonly boundarySession: { readonly sessionId: string; readonly authenticatedIso: string };
}Both files describe a constructed illustration and are written to be read, not deployed. The settlement-authority schedule shown is illustrative of an industry-standard structure; no carrier's actual schedule, limits or role names are depicted.
The limits of the argument, and what would falsify it
Four things could be wrong here, and they deserve their strongest forms rather than the weakened ones that are easy to answer.
The claim is about a prevailing pattern, not every deployment. I am describing how agent stacks propagate authority in the frameworks and reference architectures whose documentation is public. A carrier that mints a per-settlement credential bounded in dollars and naming the human authority holder already has the object, and this piece is inapplicable to it. I have not found such a deployment described in any primary source, but I have not surveyed the sector, and absence of publication is not absence. One public, checkable counterexample — a production claims chain where each hop carries a distinct, monetarily bounded grant naming its parent — would confine this piece to a description of what the rest of the field is doing.
The Evaluation Tool could arrive weakened, late, or not at all. Adoption at the Fall 2026 National Meeting is an expectation, and pilots exist to change instruments. The Tool could be adopted without the accountability material this piece leans on, or adoption could slip into 2027, or states could adopt it unevenly the way they adopted the bulletin. What that would falsify is the timing argument — the claim that the worksheet is imminent. What it would not falsify is the structural one: settlement authority is the carrier's own control, CL 7's demonstration demand is already in force in New York, and both bind whether or not the NAIC adopts anything this fall.
The strongest falsifier is behavioural, and I cannot close it. If examiners in practice accept the AIS Program paragraph — if the assertion that the platform acts under a named executive's delegated authority is treated as sufficient, worksheet or no worksheet — then the gap has no examination consequence and this argument reduces to an aesthetic preference about evidence. I have no basis for claiming examiners reject it, and I will not invent one. What I can point at is the instruments' own grammar: a bulletin that tells insurers the documentation is examinable, a circular letter that demands demonstration and refuses vendor reliance, and a regulator community that thought the questions worth piloting as structured worksheets across twelve states before adopting them.
I cannot quantify the degradation, and I am not going to pretend otherwise. The empirical question — how attribution accuracy falls as agent-to-agent handoffs rise — has, as far as I can find, no published measurement in an insurance context or any other. There is no curve in this piece and no threshold. Anyone who attaches a number to it should be asked for the primary source.
One more limit, because it cuts against the way arguments like this are usually deployed commercially. Nothing above shows that a carrier should stop building claims chains, and I am not aware of a published insurance incident turning on this failure — if I were, it would be in the sources rather than in a paragraph like this one. The argument is about what a carrier can demonstrate when examined, which is a narrower and better-behaved claim than a prediction of loss.
What an answer would have to be
This is the teardown. Building the answer here would collapse a pair of pieces into a worse single one, so I will name the properties and stop.
- The grant is an object created at the handoff, not a condition inherited from the environment and not a field stamped on afterwards. Something is constructed at the moment authority passes, or there is nothing to evidence.
- It is bounded in the decision's own units and narrows as it descends. For a claims chain that means money: a settlement agent holding a grant capped at the estimate for this claim, derived from a broader grant, derived ultimately from a rung of the authority schedule that names a person. A grant identical to its parent is the same grant and carries no information.
- It is verified where the payment lands, not only where it was issued — the payments endpoint evaluates the chain, not merely the bearer.
- It terminates in a natural person who holds the corresponding human authority, with the record of that person's decision bound into the grant rather than sitting in an adjacent document. The regime does not want a well-formed chain of services; it wants the chain to end in someone answerable, at a rung of a schedule that already exists.
And insurance adds a fifth property the banking version of this argument did not need: the record has to land where this sector's examiners already look. A grant receipt that lives only in the engineering trace has failed half its purpose; it must project into the claim file — the native evidence object — and into whatever shape the examination worksheet takes, without re-keying and without prose. That construction, including the state-variance and ECDIS-lineage machinery it needs, is the subject of the companion to this piece, "An authority record the examination worksheet can score."
The reason to build it now rather than after the Fall meeting is the pilot itself. Twelve states are spending 2026 teaching examiners to ask these questions from a script, and the script is being revised — four versions by June — against what examiners find. Whatever is adopted will have been sharpened against the records carriers can currently produce. That interval is open, and it is the only period in which what a carrier builds shapes what it is later scored against.