Start with the deviation, because the argument is not abstract and the abstraction is what usually gets it dismissed.
A granulation step runs slightly outside its established parameters at 02:41. An operator raises the deviation, by name, at the terminal. From there the work is done by agents. An investigation agent assembles the timeline from the equipment logs and the batch record. A retrieval agent pulls the product's prior deviations, the validated ranges, the relevant standard operating procedures and the last three batches made on the same line. A drafting agent writes the impact assessment, concludes that product quality is not affected, and proposes a disposition. The disposition is queued for approval. At 09:12 the next morning, a reviewer in the quality unit opens the record, reads it, and applies an electronic signature whose meaning field says approval.
The system is validated. The audit trail is complete. Every entry is timestamped, attributable to a unique account, and nothing was overwritten without the prior value surviving. If an inspector asked to see the record, it would be produced in minutes and it would be in good order.
Then someone asks a question that sounds administrative and is not:
Who decided this batch was unaffected?
The record answers a neighbouring question with total confidence. The disposition text was authored under the account svc-qms-agents-prod. The account is provisioned, owned, periodically reviewed, and restricted to the endpoints the workflow needs. The signature at 09:12 belongs to a named reviewer, is bound to the artefact, carries a timestamp, and states its meaning. Every element of that is real, evidenced and testable.
None of it answers the question. It establishes that a service produced a conclusion and that a person signed the artefact containing it. It does not establish that the person reached the conclusion, or examined what the conclusion rests on, or would have reached a different one had the chain surfaced a different set of prior batches. Attribution of an action to an account is not attribution of a judgement to a person, and the distance between them is not a documentation gap that better record-keeping closes.
The overnight granulation deviation, the account name, the timings and the reviewer are a constructed illustration, assembled from patterns documented in published regulation, public inspection literature and vendor documentation. It is not a report of any real deviation, batch, product, site, company or deployment, and no part of this piece describes client work.
The objection, stated properly
There is one strong response to what I have just written, and it comes from the people who know this estate best. It deserves to be put at its full strength, because it is a better objection than the equivalent one in most other sectors, and answering a weakened version of it would be a waste of everybody's time.
It goes like this. Pharmaceutical manufacturing did not discover computerised records last year. This industry has spent three decades building the most exacting record-keeping regime in commercial practice, and it built it specifically because records are the product of a quality system in a way they are not anywhere else. Electronic records and electronic signatures are governed by their own regulation. Computerised systems are qualified and validated before use, and revalidated after change. Data integrity expectations are explicit and are examined against directly: records must be attributable, legible, contemporaneous, original and accurate, and inspectors have been finding against the attributable limb of that for a decade. Audit trails are computer-generated, time-stamped, independent of the operator, and must not obscure previously recorded information. Signatures carry the printed name of the signer, the date and time of signing, and the meaning of the signing. Signatures are linked to their records so they cannot be excised, copied or transferred. Access is unique to individuals, and accounts are not shared.
So the natural reading of my complaint is that I have described a site with weak computerised-system controls, which is a site with a bigger problem than agents. Validate the system, apply the controls that already exist, and the record already names a person at the point where the decision is recorded.
I want to concede that at full strength before answering it, because it is right about a great deal and the thing it is right about is not the thing at issue.
The controls are real, they are mature, and they work at what they were built for. This is not a sector where the audit trail is aspirational. It is examined. Firms have received findings for audit trails that were disabled, for shared logins, for entries made after the fact and backdated, for the absence of a second-person review of the trail itself. Those controls have teeth and they close the failure modes they were designed against — the ones where somebody changed a number, or made an entry under someone else's name, or made it three days later and wrote yesterday's date. Every one of those failure modes is about the integrity of an action. All of them are now hard to commit and easy to detect, which is a genuine achievement and not one I am going to talk past.
But the unit of those controls is the entry, and the unit the regime cares about here is the decision. An audit trail records that at 04:07 the account svc-qms-agents-prod created a record whose impact-assessment field contained a particular block of text. That is an entry, and the trail is perfect about it. What the regulation is reaching for at the disposition point is different in kind: a determination, made by the unit to which the responsibility and authority to approve or reject is assigned, that this material meets its specifications. The trail can tell you when the determination was recorded and by which account. It cannot tell you that a determination occurred, because a determination is not an entry and nothing in the system was ever asked to emit one.
And the signature at the end is a countersignature, which is a weaker object than it looks. The meaning field on the signature says approval, and the signature is bound to the artefact. But the artefact was produced by a chain the signer did not run, from an evidence set the signer did not assemble, containing a conclusion the signer did not reach. What the signature establishes is that a named person, at a named time, endorsed a document. That is a real fact and it is worth having. It is not the same fact as the person having made the judgement the document asserts, and the gap between the two grows with exactly the thing an agent chain is deployed to increase — the volume of finished reasoning arriving at a reviewer's queue.
Which is the whole argument, compressed. The existing controls attribute an action to an account. The regime, at these specific points, attributes a decision to a person. An agent chain lives entirely inside the first category and produces artefacts that are consumed as though they belonged to the second.
Why these particular decisions are personal by construction
It is worth being precise about which decisions this argument reaches, because the reason it reaches them is not a general claim about accountability. It is a claim about where the regulation puts a person.
United States current good manufacturing practice regulation for finished pharmaceuticals establishes a quality control unit and assigns it responsibility and authority — the two words appear together — to approve or reject components, drug product containers, closures, in-process materials, packaging material, labelling, and drug products. It extends that to product manufactured, processed, packed or held under contract by another company, so the responsibility does not travel outward with the work. It requires the unit's responsibilities and procedures to be in writing and to be followed.
Separately, the regulation requires that production and control records be reviewed and approved by the quality control unit to determine compliance with all established, approved written procedures before a batch is released or distributed. And it requires that any unexplained discrepancy, or the failure of a batch or any of its components to meet any of its specifications, be thoroughly investigated — whether or not the batch has already been distributed — with the investigation extending to other batches of the same drug product and other drug products that may have been associated with the specific failure or discrepancy.
Read the shape of that rather than the subject matter. Three properties fall out, and all three are load-bearing for what follows.
- The holder is a named unit, not a process. The regulation does not say that batches shall be released in accordance with a validated procedure. It says a unit has the responsibility and the authority to approve or reject. A process cannot hold an authority; only a person or a body of persons can be assigned one.
- The review is prior and substantive. Records are reviewed and approved to determine compliance, before release. That is a determination made against evidence, not an acknowledgement that a workflow completed.
- The obligation propagates outward and does not attenuate. Work performed under contract by another company remains within the unit's approve-or-reject responsibility. Handing the work out does not hand the decision out.
European and United Kingdom practice puts a second person on top of this — a named qualified person who certifies each batch personally before release, with the name in a register. That is the only time I will mention it, because the argument does not need it: the United States construction already places the decision in a unit rather than in a workflow, and every market this piece is written for has an equivalent.
The physical fact: nothing was delegated, because delegation was never an object
Everything above is context. This section is the argument, and it is a claim about what exists in memory rather than a claim about what anyone failed to log.
In the prevailing propagation pattern, authority moves between agents as ambient context rather than as an event. Three mechanisms account for nearly all of it, and the property they share is that none of them produces a record because none of them is a transaction.
- A shared session. The chain runs inside one session authenticated once at the boundary. Each agent reads from it. Nothing is passed, because the session is simply in scope, so there is no moment at which the retrieval agent gives the drafting agent anything and therefore no moment at which a grant could have been recorded.
- An inherited credential. A service account, a token or a pre-constructed client sits in the process environment, built at startup. The drafting agent does not obtain it. It has held it since the process began, on identical terms to every other agent in the process.
- A prompt and a context window carried forward. Instructions, tool definitions and accumulated state cross the handoff. What changes is what the receiving agent knows. What does not change is what it may do, because capability was never modelled as the thing that travels.
Now compare that with how every other control in a quality system produces its evidence. A batch record entry is an event. A second-person verification is an event. A change-control approval is an event with a requester, an assessor and an approver. A deviation closure is an event. The entire apparatus is built out of request-perform-verify triples, and a triple leaves a trace by construction — which is precisely why the audit trail works so well against the failure modes it was designed for.
The agent handoff produces no request, no verification and therefore nothing to record. It is not that the record is thin. It is that the shape of the thing this control framework knows how to evidence was never instantiated.
This is why it is not a missing field. A missing field implies there was a value and nobody wrote it down. Here there is no value. When the drafting agent writes an impact assessment under a credential minted before the chain started, the only account identifier correct to attach to that entry is the one already attached to every other entry in the chain — which therefore distinguishes nothing. Writing it four times does not produce four facts. It produces one fact copied four times, and the copies say nothing about what happened in between.
And it is why more validation does not reach it. Qualify the infrastructure. Validate the application against a full specification. Run the operational qualification, the performance qualification and the periodic review. Every one of those establishes that the system does what it was specified to do, reproducibly. None of them creates a decision record, because the specification never called for one — the system was specified to produce an assessment and route it for signature, and it does that correctly. A validated system faithfully reproduces the shape of the thing it was built to be. If that shape has no decision object in it, validation confirms the absence rather than closing it.
It would be comfortable to treat this as something a software release closes. I do not think it is. The object requires a change in three places at once: a credential that can be derived in a weaker form without a round trip to the issuer; a handoff interface that takes the derived credential as a parameter rather than only taking context; and a check at the system of record that evaluates the chain rather than only the bearer. Each of those exists somewhere in the literature. None is the default in the agent frameworks and reference architectures whose documentation I have been able to read, and patching one of the three produces nothing usable. I have not surveyed what individual manufacturers have built privately, and a counterexample would be a fact about one deployment rather than about the field.
What the regulatory position actually is, and what it is not
There are two verified artificial-intelligence instruments from the United States drug regulator that a manufacturer would reasonably reach for here, and it is worth being exact about what each addresses, because both are routinely stretched past their scope in vendor material.
The first is a draft guidance published in January 2025, "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products." Its subject is what it says: artificial intelligence used to produce information offered in support of a regulatory decision about a drug's safety, effectiveness or quality. That is a submissions-facing question. It is a serious document and it is not a control specification for an agent chain operating inside a validated manufacturing estate.
The second is "Guiding Principles of Good AI Practice in Drug Development," published in January 2026. It is what its title says — principles, pitched at development. Principles are useful. They are not the artefact a quality unit needs when it is deciding whether an agent may draft a deviation investigation and what the record has to show afterwards.
The third thing is an absence, and it is the one that matters. As of the date on this piece, I have not been able to verify any published instrument from the drug regulator addressing agentic artificial intelligence operating inside good manufacturing practice — no control expectations for agent chains in a validated environment, no attribution standard for a decision that regulation assigns to the quality unit. I am stating that as a verified absence at a date rather than as a prediction, and it is the kind of claim that should be re-checked rather than inherited.
Readers who follow financial-services supervision will recognise the structure. Prudential supervisors in the United States revised their interagency model risk management guidance on 17 April 2026 and placed generative and agentic models outside its scope in a footnote of the shared interagency document, while returning to the institution the determination of appropriate governance and controls for what the document does not cover. I am citing that only as a structural parallel and not as law reaching a manufacturing site. The pattern is the same in both places: the specification is deferred, the obligation is not.
But the difference between the two cases runs in this sector's disfavour, and it is worth stating plainly because the comfortable reading gets it backwards.
In banking, what deferred was guidance sitting on top of guidance. Supervisory guidance in United States banking has never carried the force of law. When the agencies put agentic systems outside a guidance framework, what came out of scope was a document that already did not impose requirements on its own. The obligations survived because they lived in statute, in rules, and in the general expectation of safe and sound operation — elsewhere, and untouched.
In manufacturing, what is missing is guidance sitting on top of codified regulation. The approve-or-reject responsibility of the quality control unit is not a supervisory expectation that a document could soften. It is regulation. The requirement that production and control records be reviewed and approved before release is regulation. The requirement that a discrepancy be thoroughly investigated and the investigation extended to associated batches is regulation. None of that deferred, none of it narrowed, and none of it acquired an agent-shaped exception. What is absent is only the layer that would have told a manufacturer how to satisfy it while running agents — which means every manufacturer is currently deciding that for itself, and the eventual specification will be written against whatever the field has already built.
That is a stronger argument for building the attribution layer now, not a weaker one. The interval in which a company's choices influence what it is later measured against is open, and it is the only such interval there will be.
What fails, at mechanism level
Generic arguments about agent risk are cheap. Here is what specifically breaks, in the terms this sector uses.
Disposition, where the recommendation arrives finished. The quality unit's decision is meant to be a determination against evidence. What arrives is a completed assessment with a recommendation, produced by a chain whose reasoning is not in the record and whose evidence set is not bound to the artefact. The reviewer can agree or disagree, but disagreeing requires reconstructing work they cannot see. The practical result is that the decision migrates upstream into the chain while the accountability stays downstream with the signature, and nothing in the record marks that it moved.
Extending an investigation to other batches. This is the requirement that hurts most, and it is the one the architecture is worst at. When a discrepancy is found, the investigation must extend to other batches of the same product and other products that may have been associated with it. Scoping that population means enumerating what was affected by the defective thing. If the defective thing is a retrieval step that consistently missed a class of prior deviations, the population is every disposition that step touched — and if authority is ambient, the only boundary the record supports is everything the service account did in the window. That is dramatically larger than the true population, so the manufacturer either over-scopes at enormous cost or under-scopes on a judgement it cannot evidence.
Change control, where the approval is the whole point. A change to a validated system or an established procedure passes through an assessment and an approval. When the assessment is drafted by an agent and the approval is a signature on the draft, the change-control record has the same defect as the deviation record, with one addition: change control is the mechanism by which the estate governs itself, so a defect here propagates into every subsequent qualification decision that relies on it.
Revocation, which the estate cannot do at the right granularity. Access can be removed from a service account in seconds, and any site can demonstrate that. What cannot be done is withdrawing authority from a chain, because there is no chain object to withdraw from — only an account, and disabling the account stops everything it does, including all the work that was fine. The operational answer to a defect is therefore a blunt stoppage, in an environment where stopping the quality system stops release.
There is a fifth consequence that is less about failure and more about how failure is found. Data integrity findings in this sector have historically turned on the attributable limb: shared accounts, disabled trails, entries that could not be tied to the person who made them. An agent chain does not trip any of those. Its entries are attributable, contemporaneous and unaltered. It presents, to the exact controls this industry built, as clean. That is not a reason to be comforted. It is a reason to notice that the detection apparatus is pointed at a failure mode this one does not exhibit.
The question, asked against the record that exists
Three files, in the order a quality reviewer would meet them. The first transcribes the shape of the record an agent-assisted deviation workflow actually produces. The second attempts the disposition question against it, and the type system makes the honest answer unavoidable: the return type has no case for a person's judgement, because no input can produce one. The third is a contrast — the same query against a record built so that it terminates. Nothing here is a proposed design; the third file is what the companion piece constructs, shown only as far as is needed to see where the first two stop.
Written from the shape of a conventional agent trace plus the audit-trail and signature record a validated system produces alongside it. Note what carries across steps and what does not: the deviation identifier does, the operation does, the account does — but the account is the same value at every step, which is why it is typed as a property of the run rather than of the step.
/** A service identity. Provisioned, owned, periodically reviewed — and long-lived. */
export interface ServiceAccount {
readonly kind: "service";
readonly id: string; // e.g. "svc-qms-agents-prod"
readonly ownerOfRecord: string; // an owner is not a decider
readonly lastAccessReviewIso: string;
}
/** A person, as the quality system understands the term. */
export interface QualityUnitMember {
readonly kind: "person";
readonly directoryId: string;
readonly printedName: string;
readonly unit: "quality";
}
/** One step in the chain. Every field here observes something that happened. */
export interface WorkflowStep {
readonly stepId: string;
readonly parentStepId: string | null;
readonly operation: "invoke_agent" | "retrieve" | "draft" | "queue_for_approval";
readonly component: string; // "investigation" | "retrieval" | "drafting"
readonly startedIso: string;
/** What the step wrote into the record, if anything. */
readonly entry?: { readonly field: string; readonly bytes: number };
}
/**
* An electronic signature as the regulation shapes it: printed name, date and time,
* and the MEANING of the signing. All three are present. None of them is evidence
* that the signer performed the reasoning in the artefact they signed.
*/
export interface ElectronicSignature {
readonly signer: QualityUnitMember;
readonly signedIso: string;
readonly meaning: "review" | "approval" | "responsibility" | "authorship";
/** Bound to the artefact, so it cannot be excised or transplanted. */
readonly artefactDigest: string;
}
export interface DeviationRecord {
readonly deviationId: string;
/** The account sits here, not on WorkflowStep: every step ran under the identical
value. Modelling it per-step would let a reader believe four facts were recorded
where one was. */
readonly account: ServiceAccount;
readonly steps: readonly WorkflowStep[];
readonly disposition: { readonly text: string; readonly proposedBy: "chain" };
readonly signature: ElectronicSignature;
}All three files describe a constructed illustration and are written to be read, not deployed. The second file compiles, and that is the point: everything it returns is a fact the record genuinely supports. What it cannot return is the one thing the regulation asks for, and the closed union is what makes that visible rather than arguable.
The limits of the argument, and what would falsify it
Four things could be wrong here, and it is worth naming them at their strongest rather than in the weakened forms that are easy to answer.
The claim is about a prevailing pattern, not about every deployment. I am describing how agent stacks propagate authority in the frameworks and reference architectures currently in use. A manufacturer that mints a per-decision grant from its own authority service, bounded to the deviation at hand, naming its parent, and refusing the terminal capability by construction, already has the object and this argument does not apply to it. I have not found such a deployment described in any primary source, but I have not surveyed the sector and absence of publication is not absence. One public, checkable counterexample would confine this piece to a description of what the rest of the field is doing.
The countersignature may be doing more work than I credit. The strongest version of the objection I have not fully answered is that a competent reviewer does not merely endorse — they re-derive. If the quality unit genuinely re-performs the assessment against the underlying record, then the signature is not a countersignature and the decision really is the person's. I cannot measure how often that happens, and I am not going to assert a number. What I can say is that the economic case for the agent chain is throughput, that throughput and re-derivation pull against each other, and that the record cannot distinguish the reviewer who re-derived from the reviewer who did not — which is itself the gap, restated.
The regulator could publish, and specify exactly this. The drug regulator may issue guidance addressing agentic systems in good manufacturing practice that specifies the controls an attribution layer would provide. In that world a manufacturer that built the evidence layer early merely built it early, which is the cheapest of all the ways to be wrong. The asymmetry is the point: building early is a design constraint absorbed while the estate is small; building late means retrofitting attribution into chains already operating on commercial product.
The strongest falsifier is behavioural, and I cannot close it. If inspectors in practice accept the signature as establishing the decision — if the endorsement of an agent-produced artefact is treated as satisfying the quality unit's approve-or-reject responsibility — then the gap has no regulatory consequence and this argument reduces to an aesthetic preference about evidence. I have no basis for claiming that inspectors reject it, and I am not going to invent one. What I can point at is the text assigning responsibility and authority to a unit rather than to a workflow, the requirement that records be reviewed and approved to determine compliance before release, and the fact that the attributable limb of data integrity has been the most-cited limb for a decade in a sector that has not yet been asked this particular question.
There is one more limit worth stating, because it cuts against the way this argument is usually deployed commercially. Nothing above shows that a manufacturer should stop building agent chains, and nothing above suggests that the absence of an authority object has caused a product-quality failure anywhere. I am not aware of a published finding in this sector turning on this mechanism, and if I were it would be in the sources rather than in a paragraph like this one. The argument is about what a company can demonstrate when asked, which is a narrower claim than an argument about what will go wrong.
What an answer would have to be
This is the teardown. Building the answer here would collapse a pair of pieces into a worse single one, so I will name the properties and stop.
An authority record that survives a quality walkthrough has to have at least four properties, and the fourth is the one that usually gets dropped.
- The grant is an object created at the handoff. Not a condition inherited from the environment, and not a field stamped onto a trace afterwards. Something is constructed at the moment authority passes, or there is nothing to evidence.
- It names its parent and permits strictly less than its parent. A grant identical to its parent carries no information — it is the same grant. If the chain does not narrow, the chain is decorative.
- It binds the evidence the decision was taken against. A signature pinned to an artefact establishes that the artefact did not change. What is needed additionally is that the evidence set behind the artefact is pinned too, so that the question of what the decider could see has an answer years later rather than a reconstruction.
- The terminal capability is never granted to a chain. Draft, retrieve, assess and recommend can all be delegated. Approve, reject and release cannot, because the regulation assigns those to a unit of persons. That is not a configuration choice to be made per site; it is a capability line that has to be enforced where the action lands, so that no chain — however well-formed its grants — can cross it.
And it has to survive two conditions that a clean-room design tends to assume away: some steps run in systems the manufacturer does not own, so the chain has to be verifiable across a boundary; and the record has to remain reconstructable for as long as the product's records must be retained, which is measured against the batch rather than against the calendar and outstrips the lifetime of most of the software that produced it.
None of that is exotic. The credential constructions supporting derivation without a round trip have been in the literature for over a decade, and the regulatory structure assigning a decision to a named unit has been in force for far longer. What has not been done is the assembly, in a form a site can operate and an inspector can walk. That construction is the subject of the companion to this piece, "Attributable authority for GxP agent chains."