An insurer deploys two agents in the same quarter, and both deployments are sensible. The first sits in claims: it assembles first-notice-of-loss files, pulls the claimant's prior loss history, retrieves medical records on injury claims through a records vendor, and drafts an adjudication summary for the adjuster. The second sits in personal-lines underwriting: it triages incoming quote applications, verifies the declared facts against policy administration and loss history, fetches the external data the rating plan uses, and routes the clean cases to straight-through pricing. Each has a one-sentence job description. Each produces output a human can check. These are exactly the workflows an insurer should automate first.
Underneath them sits a set of connectors, and every one of them was approved. The policy administration connector went through integration review as a servicing tool. The claims-history connector — which fronts both the insurer's own claims system and a contributory loss-history database — was approved for adjudication support. The medical-records connector, fronting a retrieval vendor, was approved for injury claims, because that is where medical evidence is legitimate and routine. The external-data connector was approved when procurement signed the vendor contract: credit-derived scores, public-records attributes, the feeds the rating plan names. Four connectors, four tickets, four competent reviews.
Then the platform team assembled the agents, the way every agent platform currently assembles them: by attaching connectors to a toolbelt through a configuration screen. The claims agent got the connectors a claims workflow needs. The pricing agent got the connectors a pricing workflow needs. And because the medical-records connector and the loss-history connector were already registered in the same catalogue, and because the retrieval sub-agent that wraps them is shared infrastructure, the pricing agent's process can — not does, can — reach the same records the claims agent reads. For the duration of any quote it triages, its reachable set is the union of everything the four connectors reach.
Everything in this trace is a constructed illustration: the insurer, the two agents, the connectors, the tickets, the shared retrieval sub-agent and the configuration screen are invented. It is not a client engagement, not a disclosed incident, and not a description of any named carrier, vendor or platform. I have assembled it out of mechanisms that agent platforms document in their own product literature and out of duties that appear in the sector's own supervisory instruments, because the argument needs a concrete trace and I will not manufacture one out of somebody's market-conduct finding.
Now put the question in the mouth of the person who will actually ask it. A market-conduct examiner, working through an insurer's AI governance in 2027, asks: why could the pricing agent read that? Not did it — could it. The estate has an answer, and the answer is a configuration: the connector was granted in March, the toolbelt was assembled in May, here is the ticket, here is the approval. What the examiner is asking for is a decision: who decided that a pricing task could reach medical records and claims-history rows, on what basis, bounded how, reviewed when. Nobody made that decision. It emerged, from four decisions that were each about something else.
The two objections that have to be answered before anything else
There are two serious responses to this trace, and both deserve to be stated at full strength, because anyone who has worked in underwriting or in insurance compliance is already holding at least one of them.
The first: underwriting has always been broad, and the breadth is lawful. An underwriting file has always contained claims history, and on many lines it lawfully contains medical evidence: life and disability underwriting run on attending-physician statements, and casualty claims cannot be adjudicated without treatment records. Credit-derived insurance scores are filed rating factors in most US states. The loss-history databases are contributory precisely because the industry decided, decades ago, that prior claims are legitimate underwriting information. The rating plan itself is on file with the regulator, factor by factor. This is not a sector that acquired broad data access by stealth; it acquired it in public, through filings, and the filings were approved. An argument that treats breadth itself as the defect has misunderstood the business.
The second: the governance already exists, and it is newer and sharper than you are implying. The NAIC Model Bulletin on the use of AI systems by insurers, adopted in December 2023 and in force through more than twenty adopting jurisdictions by mid-2026, expects a written programme for the responsible use of AI systems: governance, risk management, internal controls, an inventory, due diligence over third-party AI systems and data, and documentation available to the regulator on request. New York went further and earlier than most: DFS Insurance Circular Letter No. 7 of 11 July 2024 addresses external consumer data and information sources and artificial intelligence in underwriting and pricing directly, and its expectations reach fairness analysis, board-level governance, third-party accountability and disclosure. Insurers have stood up committees, inventories and model-governance functions to meet exactly these instruments. This is not an ungoverned sector discovering governance; it is arguably the sector with the most specific AI supervisory material in US financial regulation.
Both objections are right, and neither will be softened here. The breadth is the business, and the governance is real. The argument survives both, and seeing why requires looking at where the duties in those instruments actually attach.
Take the breadth objection first. A filed rating plan is a bound on the decision, not a bound on the read. It names the factors that may enter a price, and in a rate-regulated line it is enforceable factor by factor. But the human system it was designed for carried an implicit second control: the underwriter who assembled the file was a person, accountable under conduct rules, whose access left an application trail and whose judgement could be deposed. The rating plan bounded what could lawfully enter the price; the person bounded, imperfectly but accountably, what actually reached the decision. An agent estate keeps the first control and silently deletes the second. The reachable set of a pricing agent is now a property of a credential composition no filing ever described, and the only artefact bounding what could have entered the decision is a configuration nobody reviewed as a whole.
And the direction of the duty makes that gap operational rather than philosophical. When the outcome of an underwriting or pricing decision is adverse, the sector's instruments expect the insurer to state the reason — specifically, including where the decisive information came from, not as a gesture at a model. An unfair-discrimination analysis, of the kind New York's circular letter expects insurers to perform before using external data or AI at all, is an analysis of what entered the decision. Both of those are claims about a set: the inputs. A union-scoped agent can show what it was told to use — the prompt, the rating plan, the instructions. It cannot show what it could not have used, because the credential admitted everything and no record was ever made of the difference. The insurer is left asserting the negative — the medical file did not enter the price — from a system whose logs establish only that the door was open.
Now the governance objection. A written AI programme of the kind the Model Bulletin expects is an inventory of systems and models, a lifecycle, a testing regime, a set of committee minutes. Its unit of account is the model or the system. The examiner's questions, when they come, attach to decisions: this declination, this rate, this claim. Between the programme and the run there is currently no object — nothing that says this task, on this date, for this purpose, was entitled to reach these sources and no others. The governance is real and it is aimed one level too high. It governs the things that were built. It does not govern the compositions they form at runtime, because nothing in the estate reifies a composition long enough for anyone to govern it.
A name is not a bound, and a vendor feed is not even a name
The physical fact underneath the trace is the same one that recurs across every sector this series has examined: a tool grant is written as a name and exercised as a transitive closure. A connector, at the level a runtime sees it, is a name, a description, an argument schema and a credential. Nothing in that list constrains which policies, claimants, claims or data elements the tool may reach when invoked. The argument schema types the arguments; the reachable set is a property of the credential; and the credential was provisioned against an integration, not a task.
Insurance adds a twist that makes the closure worse than in most sectors, and it deserves its own paragraph. One of the connectors in the trace fronts an external data vendor. The reachable set of that connector is not a set of rows in a system the insurer owns; it is whatever the vendor's product returns. When the vendor enriches the feed — adds a lifestyle attribute, a new public-records linkage, a refreshed score version — the reachable set of every agent holding that connector grows, and it grows without a new approval, without a new filing, and typically without the insurer's model-governance function learning of it on any schedule better than an annual vendor review. The closure is not merely uncomputed. Part of it is defined by a third party's product roadmap.
New York's circular letter is unusually direct about where the accountability for that sits, and the direction matters: the insurer remains responsible for the external data and the tools built on it, whoever developed them. The expectation, paraphrased, is that an insurer understands the data it is using, has assessed it, and can demonstrate the assessment — and that reliance on the vendor's own assurances is not an answer. Read against the mechanism just described, that is an expectation the connector model cannot meet by construction: an insurer cannot have assessed data elements that did not exist when the connector was approved, and the connector is how the elements arrive.
Paraphrase discipline: the expectations attributed to Insurance Circular Letter No. 7 (2024) in this piece — fairness analysis before use, governance proportionate to risk, retained documentation available to the Department, third-party accountability resting with the insurer, and specific reasons for adverse decisions — are paraphrases of a short letter that should be read whole, not quotations. Where this piece needs the letter's exact words, it does not supply them, deliberately: check the text at dfs.ny.gov before relying on any characterisation, including mine.
Then there is composition, which is where the arithmetic stops being additive. The medical-records connector, approved for injury claims, is unremarkable. The pricing agent's straight-through-quote toolbelt is unremarkable. A shared retrieval sub-agent is ordinary platform engineering — nobody builds two document fetchers. But held together in one estate, they are a path: medical information, reachable by a pricing process. On life and disability lines, where medical underwriting is lawful, that path at least corresponds to a filed practice. On the personal motor line in the trace, it corresponds to nothing an actuary filed, nothing a committee approved, and nothing anyone would defend on purpose — and it exists anyway, because the unit of approval was the connector and the unit of authority is the set.
The contributory databases add a final composition edge that insurance-adjacent readers will recognise and others may not: they are contributory. What an insurer reads from a shared loss-history database, it also feeds. An agent estate wired into such a feed does not just have a read closure; part of its closure leaves the estate, into an industry utility, under contribution rules written for human-operated claims systems. A grant that looks like read a claims file quietly includes populate an industry record about a person, and no connector-level review received that pair either.
The instruments already ask decision-shaped questions
The comfortable conclusion at this point would be that the supervisory material was written for models and needs updating for agents. As in healthcare, the comfortable conclusion is false, and it is worth walking the instruments to see why: the questions they ask are already at the granularity of the decision, and agents are simply the first deployment where the estate cannot answer them.
Start with New York. Circular Letter No. 7's core move is to attach conditions to use: an insurer should not use external consumer data or AI in underwriting or pricing unless it has established, through its own comprehensive assessment, that the use does not rely on prohibited characteristics and does not produce unfairly discriminatory outcomes — and it should be able to show the analysis. Every verb in that sentence attaches to a decision process, not to an integration. An assessment of a data source is an assessment of its role in decisions. A demonstration that outcomes are not unfairly discriminatory is a demonstration about what entered the outcomes. An insurer whose agents hold union-scoped credentials can present its assessment of the sources it intended to use, and then has no artefact — none — establishing that those were the only sources in reach of the decision path. The assessment governs the menu. The credential set the table.
The transparency expectation sharpens it. When a decision is adverse, the insurer is expected to tell the consumer the actual reason, with enough specificity to be meaningful — and the letter is explicit, in paraphrase, that opacity of a vendor's model is not an excuse. Answering that expectation requires knowing, per decision, which information the decision rested on. Today that answer is assembled from the model's feature list and the workflow's design documents: from intent. An agent that could have read beyond its intent breaks the assembly, because the feature list no longer bounds the inputs; it only describes the ones that were supposed to be there.
The NAIC layer asks the same questions with an examiner's grammar. The Model Bulletin's programme expectations end, operationally, in a market-conduct examination: the regulator may request the insurer's AI system inventory, its governance records, its testing evidence, its third-party due diligence — and the books-and-records duty reaches information about systems the insurer did not build. What converts those expectations from prose into a worksheet is the AI Systems Evaluation Tool: an examiner-facing instrument, piloted by twelve states between January and September 2026, at version 4.0 as of the June 2026 public discussion, with adoption expected at the Fall 2026 National Meeting. The pilot ends next month, as this piece is published. Whatever its final text says — and it is not final, which this argument respects — its direction is not in doubt: it is a structured set of questions a state examiner asks about an insurer's AI systems, their data, their governance and their outcomes. It is the closest thing to a published examination script that exists anywhere in US financial regulation's treatment of AI.
Set the two artefacts side by side: an examiner worksheet whose questions are about decisions, data provenance and demonstrable bounds; and an estate whose only authority records are connector grants and tool-call logs. The worksheet asks what entered the decision and how the insurer knows. The log answers which connector fired and when. The worksheet asks how third-party data was assessed before use. The log holds a service-account identifier. The worksheet — in any plausible final form — asks the question this piece is named for, and the estate answers with a configuration. That mismatch, multiplied by every adopting state after Fall 2026, is the sector's version of the gap this series keeps finding: the record the instrument presupposes is one the current stack never mints.
It is worth being precise about what the banking parallel does and does not add here, because insurers keep being told to look at it. When the US federal banking agencies revised their model risk guidance in April 2026 — OCC Bulletin 2026-13, with the Federal Reserve's parallel issuance SR 26-2 — the interagency text stated that generative and agentic AI models are novel and rapidly evolving and are not within its scope, with separate guidance promised. Read as a deferral rather than an exemption, the lesson for insurance is structural: nobody handed banks an agent-control specification, and nobody is going to hand insurers one either. But insurance is ahead of banking in one concrete respect, and the difference cuts against waiting: insurance already has its examination instrument in pilot. Banks are waiting for a request for information that has not arrived; insurers can read a worksheet draft. The sector that can see the exam coming has the least excuse for building estates that cannot answer it.
India is drafting now, and the Gulf has not written the instrument
The Indian thread matters because of where its clock is, not because its text exists — it does not, and that is the point. On 18 June 2026 IRDAI constituted a working group on artificial intelligence in insurance with a three-month window to report. That window is open as this piece is published and closes around September 2026. India's insurers are deploying the same agent architectures on the same platforms as everyone else, which means the composition mechanism described above is accumulating in Indian estates during exactly the period in which India's instrument is being drafted. The play this creates is the same one the banking lane has with the RBI's unfinalised model-risk guidance: the frameworks insurers build inside the window are the installed base the working group's recommendations will be written against. An Indian insurer that can already answer the decision-shaped question — why could this agent read that, and who decided — will find the eventual instrument describing its practice. One that cannot will find the instrument describing its gap.
On the Gulf, the honest sentence is short: I could verify no insurance-specific AI instrument in any GCC jurisdiction as of this writing, and this piece therefore cites none. The region's 2026 supervisory energy is real but is running through structural and cross-sector channels rather than insurance supervision. For a carrier or reinsurer operating there, the practical consequence is that the binding expectations on an agent estate arrive through counterparties — group policies, reinsurance security committees, bancassurance partners governed as financial institutions — which tend to transmit exactly the NAIC- and DFS-shaped questions examined above. An absence of a local instrument is an absence of a local answer sheet, not an absence of the exam.
What a reviewer can measure today
As in the healthcare teardown, the honest instrument here is a thermometer rather than a thermostat: something that makes the composed reach visible without pretending to fix it. The composition audit below is deliberately small. It takes the connector grants an estate already has, at the level of description an integration review already produces, and reports the properties of the sets that toolbelts assemble from them — including the two findings an insurance examiner would ask about first: a medical-to-pricing path, and an external-data source reachable by a decisioning task without an assessment reference attached anywhere.
A thermometer, not a thermostat
A measuring instrument, not the remedy. It narrows nothing, mints nothing and enforces nothing — the construction that does is the companion piece. It exists to turn a set of connector grants and toolbelt assignments into sentences a reviewer has to look at, because the argument of this piece is that those sentences currently go unwritten, not that writing them is hard. No dependencies, no network calls.
The important design choice is that the audit runs over the composition, not the connector: every finding below is a property of a set that no single approval ever saw. Note that a vendor-controlled connector is reported as open-ended rather than expanded into today's element list — a snapshot of a set the vendor can grow is not a bound, and reporting it as one would concede the argument before making it.
/** The data domains an insurance integration review already names. */
export type DataDomain =
| "policy-admin"
| "claims-history"
| "contributory-loss-database"
| "medical-records"
| "ecdis"
| "payments"
| "outbound-communications";
/** What a decisioning workflow is for. Typed, because the findings depend on it. */
export type WorkflowPurpose =
| "pricing"
| "underwriting-eligibility"
| "claims-adjudication"
| "fraud-investigation"
| "servicing";
export interface ConnectorGrant {
readonly id: string;
readonly name: string;
readonly domains: readonly DataDomain[];
/** True where the reachable elements are set by a third party's product,
* not by the insurer's own schema — the ECDIS case. */
readonly vendorControlled: boolean;
/** Reference to the assessment that cleared this source for decisioning
* use (the CL7-style comprehensive assessment), if one exists. */
readonly assessmentRef: string | null;
/** True where reads also contribute outward, e.g. a contributory database. */
readonly contributesOutward: boolean;
}
export interface Toolbelt {
readonly agent: string;
readonly purpose: WorkflowPurpose;
readonly connectorIds: readonly string[];
}
export interface CompositionFinding {
readonly agent: string;
readonly severity: "flag" | "note";
readonly finding: string;
}
export interface CompositionReport {
readonly agent: string;
readonly purpose: WorkflowPurpose;
readonly composedDomains: readonly DataDomain[];
readonly openEnded: boolean;
readonly findings: readonly CompositionFinding[];
}
const DECISIONING: readonly WorkflowPurpose[] = [
"pricing",
"underwriting-eligibility",
"claims-adjudication",
];
export function auditComposition(
grants: readonly ConnectorGrant[],
toolbelt: Toolbelt,
): CompositionReport {
const held = toolbelt.connectorIds
.map((id) => grants.find((g) => g.id === id))
.filter((g): g is ConnectorGrant => g !== undefined);
const composedDomains = [...new Set(held.flatMap((g) => g.domains))];
const openEnded = held.some((g) => g.vendorControlled);
const findings: CompositionFinding[] = [];
const isDecisioning = DECISIONING.includes(toolbelt.purpose);
if (
isDecisioning &&
toolbelt.purpose !== "claims-adjudication" &&
composedDomains.includes("medical-records")
) {
findings.push({
agent: toolbelt.agent,
severity: "flag",
finding:
"Medical records are reachable by a " +
toolbelt.purpose +
" task. No single connector approval received this pair as its input.",
});
}
for (const grant of held) {
if (isDecisioning && grant.domains.includes("ecdis") && grant.assessmentRef === null) {
findings.push({
agent: toolbelt.agent,
severity: "flag",
finding:
"External data source '" +
grant.name +
"' is reachable by a decisioning task with no assessment reference attached anywhere in the grant.",
});
}
if (grant.vendorControlled) {
findings.push({
agent: toolbelt.agent,
severity: "note",
finding:
"Connector '" +
grant.name +
"' is vendor-controlled: its reachable elements are open-ended and grow without a new approval. Reported as open-ended, not as today's element list.",
});
}
if (grant.contributesOutward) {
findings.push({
agent: toolbelt.agent,
severity: "note",
finding:
"Connector '" +
grant.name +
"' contributes outward: part of this composition's closure leaves the estate into an industry record.",
});
}
}
return {
agent: toolbelt.agent,
purpose: toolbelt.purpose,
composedDomains,
openEnded,
findings,
};
}Nothing here decides anything, and the domain labels are the coarse ones an integration review already uses — deliberately, because the audit's value is that it runs against artefacts an insurer already has. If running it against a real estate produces zero findings, part of this essay's argument is falsified for that estate, and that is exactly the check a reviewer should run before accepting the argument.
Where this argument stops, and what would falsify it
The honest limits, stated rather than found.
The size of the composed gap is unmeasured, and I will not guess at it. No regulator, standards body or vendor publishes a measurement of the ratio between what an insurance agent's task requires and what its composed connector grants reach. The mechanism is documented — connector-level scoping is how every major agent platform provisions access, and vendor-controlled feeds are open-ended by their commercial nature — but the consequence is unquantified, and any multiplier offered in either direction should be treated as unsourced. The claim here is narrower: the composition is not merely large, it is unreviewed and uncomputed. That is the finding.
Filed rating plans genuinely bound the decision in rate-regulated lines. Where a line is rate-regulated and the pricing model's inputs are pinned, feature by feature, to the filed plan — and where those inputs are logged per decision — an insurer has a real, decision-shaped artefact for the price itself, and part of this argument weakens for that line. What survives even there is the reachability question: a pinned feature list bounds what the model was given, not what the agent around the model could read while assembling it. But a reader who runs pinned-input pricing with per-decision input logs should hold this essay to that distinction, because it is doing real work.
The examiner worksheet is not final, and I am arguing from its direction. The Evaluation Tool is in pilot; its version number has moved during 2026 and its final text will be settled after the pilot ends in September. Nothing above quotes it, and nothing should — the argument rests on the kind of instrument it is, an examiner-facing structured questionnaire about AI systems and their data, which is publicly established, not on any particular question, which is not. If the adopted tool confined itself to configuration-shaped questions, the sharpest paragraph of this piece would lose its edge. I do not expect that, and I have committed to the expectation in print.
Nothing here shows that harm occurred. No incident is described, because I do not have one and would not use somebody's if I did. The argument is about a question that cannot be answered from the record, not about damage that was done. In this sector the distinction is decisive in both directions: an insurer can have a defenceless record and a spotless book of decisions — and it will still be the record, not the book, that the examination reads first.
What would falsify the argument is a demonstration that the decision-shaped answer can be reconstructed after the fact from artefacts that already exist: that from connector grants, toolbelt configurations and tool-call logs, an insurer can produce, for one adverse pricing decision, a defensible statement of which data elements could and could not have entered it, with sources and assessment references, sufficient for an examiner — without having recorded any per-task grant at the time. I do not believe that reconstruction is possible, because the information was never present in the objects being logged. But it is a checkable claim, and the composition audit above is one way to begin checking it against an estate rather than against my word.
The shape of what is missing
Everything above converges on a single absent object. The grants an insurer's agents exercise are bound to integrations. The decisions those agents produce are bound to quotes and claims. Between them there is nothing: no record that says this quote-triage run, on this date, under this filed plan, was entitled to read these elements from these sources, with these assessments behind them, and nothing else — and here is the identifier tying that entitlement to the price it produced. The fields of that object are not exotic; they are the fields the sector's own instruments keep asking for. Its boundary objects — the quote, the claim — are already first-class in every insurance system, with identifiers, owners, and lifecycles that open and close. What is missing is the moment of minting, because the start of a task is not currently a moment at which any authority decision is made.
How that object gets minted; how the lineage of an external data source travels on the grant rather than in a vendor contract nobody rereads; how the boundary objects behave when a quote binds or a claim reopens; and what the record projects into when the examiner arrives — that is a construction, and it is the subject of the companion to this piece, "Per-claim and per-quote capability minting for insurance agents", published alongside it. This essay deliberately stops here. The sentence worth carrying out of it is the one the whole argument reduces to: in insurance, the read is never just a read — what an agent could reach is evidence about what could have entered a price, and until scope is granted at the size of the task, no insurer can put a boundary around that evidence.