Start with a batch, because the argument is concrete and the abstraction is what lets it be waved away.
A deviation is raised on the packaging line: a label reconciliation count is out by a quantity small enough to be a counting error and large enough that the procedure requires it to be investigated. An agent-assisted triage workflow picks it up. It retrieves the relevant procedure, the batch record, the equipment log and the last four comparable deviations. It classifies the deviation as minor. It drafts the investigation summary, populating the fields the quality management system expects. A quality reviewer reads the draft, agrees, and signs. Two weeks later the batch is dispositioned and released. The product goes to market and, over the following months, to patients.
Nothing about that is reckless, and it is important to say so before the rest of the argument. A human reviewed and signed. The classification was defensible. The system was validated before it went into use. The workflow sits inside a quality management system that has passed inspection before. If you were looking for the failure in that description, there isn't one.
The engineering evidence is also excellent, in its own terms. There is a trace. Every retrieval call is recorded with its query and its latency. The classification step records the model output and the token count. The drafted summary carries an artefact identifier. If the workflow had timed out, an engineer would find the cause in ninety seconds.
Four years later, an investigator opens the file. The product has been on the market and is either still within shelf life or recently past it; the batch-associated records are still inside the retention floor the predicate rules impose, which runs past expiry rather than past release. The investigator is not auditing the platform. The investigator is doing what investigators do: pulling a thread. The thread is this deviation, and the question is:
Which version of the procedure was in front of the person who signed this, and on what basis was it classified minor?
The record answers with total confidence a set of questions that were not asked. It establishes that a retrieval call was made against the document store at 14:07:22.431, that it returned in 240 milliseconds, that a classification step emitted the label minor, that a summary artefact was created with a given identifier, and that a named reviewer applied an electronic signature at 16:41 the following day. Every one of those facts is real, timestamped and, where it is still within its retention window, retrievable.
None of them is an answer. The query string tells you what was asked of the document store; it does not tell you which documents came back, in which effective versions, or which of them the reviewer actually read. The label tells you the output; it does not tell you the basis. The signature tells you that a person accepted a draft; it does not tell you what was in front of them when they accepted it. And by year four, most of the channels that could have been interrogated to reconstruct any of this have expired on schedules that had nothing to do with the question.
The label reconciliation deviation, the triage workflow, the classification, the timings and the four-year inspection are a constructed illustration, assembled from patterns documented in published regulatory guidance, industry practice literature and vendor deployment documentation. It is not a report of any real site, product, batch, deviation, manufacturer or deployment, and no part of this piece describes client work.
The objection, stated properly
Two strong responses arrive at this point, and they come from different people. The first comes from inside quality assurance, and it is much the better one. The second comes from the regulatory reading, and it is the one I hear more often. Both deserve to be put at full strength before either is answered, because the sector I am writing about does not lack for people who have thought about records for a living.
The response from quality assurance goes like this. Pharmaceutical manufacturing has the most developed record discipline in the private economy, and it did not arrive there by accident. Computerised system validation exists precisely to establish, with documented evidence, that a system performs as intended in its operating environment, and it is applied on a risk basis to every system that touches product quality. The sector's data-integrity expectations — that records be attributable, legible, contemporaneous, original and accurate — are trained into every operator and audited relentlessly. Electronic record and electronic signature requirements demand secure, computer-generated, time-stamped audit trails that record the operator, the action and the time, and that do not obscure previously recorded information. Change control governs what may be altered and by whom. Nobody in this sector forgot to write things down. Introducing an agent into a validated workflow means validating the agent-containing workflow, and if the validation is done properly the resulting records are, by construction, the records the regime requires.
The second response is regulatory. FDA has not been silent on artificial intelligence. In January 2025 it published a draft guidance, "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products," proposing a risk-based credibility assessment framework for AI models whose outputs support regulatory decisions. In January 2026 it published "Guiding Principles of Good AI Practice in Drug Development." On the economical reading, the agency has spoken, twice, recently; the framework exists; apply it and move on.
I want to concede both at full strength, because each is right about something, and in each case the thing it is right about is not the thing at issue.
Validation establishes that a system performs as intended. An inspection asks what a particular decision rested on. Those are different objects, and the gap between them is not closed by doing more validation. Validation is a statement about a system across its intended use: given inputs of this class, it produces outputs of this class, reproducibly, within these bounds, and here is the documented evidence. That is exactly what you want before the system goes into use. It says nothing about the twelfth deviation of the following March. An inspection is retrospective and particular: not "does this system work", but "what happened here, on this batch, and what was it based on". A perfectly validated system can leave a record that cannot answer a particular question, and the validation package will not tell you that it does, because answering particular questions was never what the package was scoped to demonstrate.
Data-integrity expectations attach to the records a system was designed to create. Attributable, legible, contemporaneous, original, accurate — read those as what they are: quality attributes of records, not a requirement that any particular record exist. A span emitted by an observability library is attributable to a service, legible to an engineer, contemporaneous to the millisecond, original and accurate. It passes. It is also not a decision record, and no amount of integrity applied to it converts it into one. The expectations are a floor under the records the system creates; they are not a specification of which records a system must create. That specification lives in the predicate rules and in the procedures, and neither was written with an agent-assisted triage step in mind.
The audit trail is a change log, not a basis record. This is the substitution that does most of the work in a control narrative, and it is worth being exact about. The audit trail requirement is about the integrity of the record over time: who created or modified an entry, what the previous value was, when, and — where the regime asks for it — why the change was made. It is a superb instrument for the question "has this record been altered, and by whom". It is silent on the question "what information was in front of the person when they made the entry in the first place", because that information was never an entry. The audit trail sits on top of the record. What is missing here is underneath it.
Which compresses the whole argument into one sentence. The sector's record apparatus is built to protect records from corruption over time, and it does that better than any other sector does. It is not built to capture the basis of a decision that a machine helped to assemble, because until recently the basis of a decision lived in a human head and was reconstructed, when necessary, by asking the human.
That reconstruction-by-interview mechanism is the thing that is quietly breaking, and it is worth naming plainly. It has always been how thin records got thickened at inspection: the investigator asks, people who were there remember, and the memory is corroborated against whatever documentary fragments exist. It works because a human decision-maker who read four documents and formed a judgement can usually say, years later, roughly what they read and roughly why. When the reading and the drafting were done by a system that retained none of it, and the human contribution was to agree with a draft, there is much less to remember, and what there is to remember is a review rather than an investigation.
What the two instruments actually face
The regulatory objection deserves the same precision, and it fails on placement rather than on existence.
The January 2025 draft guidance faces submissions. Its subject is the use of artificial intelligence to produce information or evidence that supports a regulatory decision about the safety, effectiveness or quality of a drug or biological product — a model whose output goes to the agency, directly or through the reasoning of an application. The framework it proposes is a risk-based credibility assessment: establish the question the model is being used to answer, establish the context of use, and calibrate the evidence of credibility to the risk that the model's output is wrong in that context. That is a good framework and it is aimed at a real problem. It is aimed at a model in a submission.
The January 2026 guiding principles face development. They are stated at the level of principles — the register in which a guiding-principles document is written — and they set out what good practice looks like across the drug development lifecycle. Principles are useful for orientation and they are not controls. Nobody validates against a principle.
Now place the workflows this piece is about. Deviation triage is not a submission. CAPA drafting is not a submission. Complaint categorisation is not a submission. Batch disposition support is not a submission. These are GxP quality operations, running inside the manufacturer's own quality system, and their outputs are records that an inspector reads rather than documents an agency reviews. The cell they occupy on the map has no instrument in it. That is a verified absence rather than a prediction: as of the date on this piece, FDA has issued no guidance facing the use of agentic AI in GxP manufacturing and quality operations.
Readers who follow the banking argument will recognise the shape, and the parallel is worth drawing once and then leaving alone, because it is a parallel and not an authority. On 17 April 2026 the Federal Reserve, the FDIC and the OCC issued revised interagency guidance on model risk management, distributed as the SR 26-2 attachment and as OCC Bulletin 2026-13. Its third footnote places generative and agentic AI models outside the guidance's scope — and the same footnote then returns to the institution the responsibility for determining appropriate governance and controls for anything the document does not cover. That is a deferral with the obligation left in place and the specification withdrawn. Nothing in the pharmaceutical instrument set says anything as explicit as that footnote in either direction. But the structural position is identical: the obligations attached to the underlying activity are untouched, and the framework that would have specified the controls for the new mechanism has not been written. Neither situation is a permission. Both are intervals.
One asymmetry makes this cell harder rather than easier than the banking one. Bank supervision proceeds largely by examination against a framework, and when the framework is silent the examiner has less to hold. Pharmaceutical inspection proceeds against predicate rules that are already binding and already general: the requirement that quality-affecting activities be documented, that investigations be thorough, that decisions be justified. An inspector does not need a new instrument about agents in order to ask what a disposition rested on. The question is available to them today, under rules that have been in force for decades, and the answer will be judged against those rules rather than against whatever guidance eventually arrives.
The physical fact: the telemetry is a record of calls, not of a decision
Everything above is context. This section is the claim, and it is a claim about what exists in storage rather than a complaint about diligence.
An agent-assisted workflow emits three families of artefact, and it is worth being specific about what each one is, because the families are routinely spoken of as though they were one thing called "the audit trail of the AI".
- Spans. A span is an observation of an operation that happened, with a start, a duration, a parent and a set of attributes. It exists because the code performing the operation could observe itself performing it. Spans compose into a call graph, and the call graph is genuinely complete: every edge in it is an invocation, and every invocation is an event the runtime can see. Spans are also retained on the observability system's schedule, which in most estates is measured in weeks to a few months, because the sizing was done against the question "can we debug last Tuesday".
- Tool call records. Arguments in, results out, occasionally truncated for size. This is the richest family and the one most often mistaken for evidence of reasoning. A retrieval call record tells you the query and, if you are lucky and nobody truncated it, the identifiers of what came back. It does not tell you what the model attended to, and it does not tell you the effective version of the document behind an identifier unless the document store was asked for a version and recorded the answer.
- Model outputs. A label, a draft, a summary, a confidence figure if the vendor exposes one. The output is the conclusion. The basis for the conclusion is not an artefact the system holds, because the mechanism that produced it does not decompose into one — and a generated explanation of the conclusion, produced after the fact by the same mechanism, is a second output rather than a record of the first.
Now set that against what a batch record is. A batch record is not a log of what the operators did. It is a designed document: the fields exist because someone decided in advance which facts must be captured for this product to be dispositionable, and the record is complete when those fields are filled. Every serious record in this sector has that character. The investigation report has required sections. The CAPA has an effectiveness check with a defined form. The disposition has a defined basis. These are not traces. They are structured artefacts, designed backwards from the questions that will be asked of them, which is precisely why they survive inspection years later.
This is why it is not a missing field. A missing field implies a value existed and nobody wrote it down. Add the field and the problem closes. Here the values mostly do not exist. "Which effective version of the specification informed this classification" is not a fact sitting unrecorded in the platform's memory — it is a fact that would have had to be constructed at the moment of retrieval, by asking the document store for a version identifier and binding that identifier to the reasoning step that consumed it. Nothing in the ordinary retrieval path does that, because the ordinary retrieval path was built to return relevant text quickly.
And it is why retention alone does not reach it. Suppose you turn sampling off, retain every span at full fidelity, and pay for seven years of storage. You now have a perfect, expensive record of which calls happened, and the inspection question is exactly as unanswerable as it was at ninety days. Retention preserves what was captured. It does not create what was not. The retention problem is real and I will come back to it, but it is the second problem, and estates that solve it first end up paying to store the wrong thing for a decade.
What fails, at mechanism level, across four workflows
Generic warnings about agent risk are cheap and this sector has heard plenty of them. Here is what specifically fails, in the four places agent assistance is actually landing in quality operations, stated in the terms the sector uses.
Deviation triage: the classification determines the depth of the investigation. Classifying a deviation as minor rather than major is not an administrative act; it sets how far the investigation goes, whether a root-cause methodology is applied, whether other batches are assessed, and whether the disposition can proceed on the timeline. If the classification was agent-assisted and the basis is not recorded, then an investigator years later cannot test the classification — they can only observe that a classification was made and accepted. Where a pattern of classifications is later found to have been systematically light, the absence of per-case basis records converts a question about individual judgements into a question about the whole population, which is a far worse conversation to be having.
CAPA drafting: the effectiveness argument is the part that has to be reconstructible. A corrective and preventive action is judged, at inspection, on whether the action addressed the actual cause and whether the effectiveness check was capable of detecting failure. Both of those are arguments rather than facts. When an agent drafts the CAPA — proposing the action, proposing the effectiveness check, populating the rationale — the argument enters the record as prose with no attached basis. A reviewer who agrees with a well-written argument leaves behind a signature on a well-written argument. Nothing distinguishes that from a reviewer who constructed the argument, and the distinction is exactly what an investigator probing a repeat deviation will want.
Complaint categorisation: the population is the regulated object. Complaint handling turns on whether a complaint is a potential quality defect, whether it is reportable, and whether a set of complaints constitutes a trend. Those are population judgements built from individual categorisations. Agent-assisted categorisation at volume is one of the most attractive applications in the whole quality function, precisely because the volume is high and the individual judgements are mostly routine. It is also the one where an undetected systematic bias in categorisation produces the largest downstream failure, because the trend analysis inherits the bias and reports that nothing is trending. Without per-decision basis records, the only way to find that is to re-review the population manually, which is the work the automation was bought to avoid.
Regulatory-submission assembly: this is the one cell where an instrument does point. Assembling a submission from source documents is an obvious agent workload and it is the workload closest to the January 2025 draft guidance, because the output goes to the agency. It is also where the evidence question is sharpest, because the assertion the submission makes is that the content faithfully represents the underlying data. If content was assembled by an agent, the faithfulness claim needs a basis: which source, which version, which transformation. A credibility assessment framework aimed at models producing evidence is the right family of instrument for this, and organisations should be reading it. Notice, though, what that does to the argument overall: the cell with an instrument is the cell furthest from where most of the deployment volume is.
Two operational consequences follow, and they are the ones that actually hurt when something surfaces.
The first is population scoping. When a defect is found — a mis-specified retrieval, a document store that was serving a superseded procedure for a period, a classification prompt that was changed without assessment — the immediate question is which decisions were affected. Answering it requires knowing, per decision, what the decision consumed. If that is not recorded, the defensible population is every decision the workflow touched in the window, which will be vastly larger than the true population. The manufacturer then chooses between an over-scoped review at serious cost and a narrower scope resting on a judgement they cannot evidence. That is a bad choice created by an architecture, and it is the same bad choice a bank faces when it cannot scope a customer-affecting defect.
The second is retention mismatch, and it is the consequence unique to this sector. The predicate rules keep batch-associated records past the product's expiration date, which for a typical product puts the record's life well beyond three years and often considerably longer. Observability platforms retain on a cost curve, and ninety days is a common default. Model versions are deprecated by providers on their own schedules, which are not the manufacturer's to set. Platform contracts get renegotiated or terminated. Every one of those is a reasonable decision by the party making it, and none of them was taken against the retention obligation attached to the batch. The result is an estate where the regulated record persists and everything that would explain it does not.
The limits of the argument, and what would falsify it
Four things could be wrong here, and I would rather state them at their strongest than have a reader do it for me.
The claim is about a prevailing pattern, not about every deployment. I am describing what agent platforms emit by default and what quality organisations are integrating today. A manufacturer whose retrieval layer returns effective-version identifiers, binds them to the reasoning step, and writes a structured decision record into the quality system alongside the human signature already has the object, and this entire argument is inapplicable to them. I have not found such a design described in a primary source, but I have not surveyed the sector, and absence of publication is not absence. One checkable counterexample — a validated GxP workflow where the per-decision basis is a first-class record retained on the product's clock — would confine this piece to a description of what the rest of the field is doing.
The reconstruction-by-interview mechanism may hold up better than I think. The sector has always reconstructed thin records through people, and it is possible that agent-assisted workflows leave reviewers with enough recall to sustain that, especially where the reviewer genuinely re-derived the conclusion rather than agreeing with a draft. I have no measurement of how reviewer recall degrades when the drafting was done by a system, and I am not going to invent one. If it turns out that reviewers of agent drafts remember as much as authors of their own investigations, this piece overstates the loss.
The instrument could arrive early and specify exactly this. FDA could publish guidance facing AI in GxP operations that names the decision record, the version binding and the retention alignment. In that world an organisation that built the evidence layer early merely built it early, which is the cheapest way to be wrong. The asymmetry is the argument: building it now is a design constraint absorbed while the workflows are young, and building it later is retrofitting basis records onto decisions already made — which is the same problem as reconstructing them, which is the problem.
The strongest falsifier is behavioural, and I cannot close it. If investigators in practice accept the trace plus the reviewer's signature as sufficient evidence of a decision's basis, then the gap has no inspection consequence and this reduces to a preference about record design. I have no basis for asserting that investigators reject it, and no published finding turning on this failure mode is known to me — if one were, it would be in the sources rather than in a paragraph like this. What I can point at is the character of the predicate rules, which ask for decisions to be justified rather than merely recorded, and the fact that the interval between decision and question in this sector is measured in years rather than months.
One further limit, because it cuts against how this argument is usually deployed commercially. Nothing above shows that agent assistance in quality operations is a bad idea, and nothing above suggests that the absence of decision records has caused harm anywhere. The workloads are genuinely valuable and the sector's caution about them is already high. The argument is about what a manufacturer can demonstrate when asked, four years later, which is a narrower and more defensible claim than an argument about what will go wrong.
What an answer would have to be
This is the teardown. Building the answer here would produce a worse single piece in place of a pair, so I will name the properties and stop.
An evidence layer that survives an inspection years after the fact has to have at least four properties, and the fourth is the one that gets dropped.
- The decision record is a first-class object, created at the decision. Not derived from telemetry afterwards, and not a field appended to a span. A structured artefact with the same status as an investigation report: inputs consumed with their effective versions, the basis on which the conclusion was reached, the alternatives that were considered and rejected, and the approver as a named person acting under a defined authority.
- It is held separately from the telemetry, on the product's clock. Telemetry is an engineering asset with an engineering retention policy, and it should keep one. The decision record belongs with the batch-associated records and inherits their retention floor, which runs past the product's expiration date rather than past the quarter.
- Version binding happens at retrieval, not at review. The effective version of a procedure or specification has to be captured when the document is consumed. Recovering it later from a document management system's history is a reconstruction, and reconstructions are what this whole argument says do not survive.
- Reconstruction is drilled as a periodic control, not assumed. Every property above can be satisfied on paper and fail in practice. The only way to know whether the record answers the question is to ask the question of a real past decision, on a schedule, and record what could and could not be answered. A quality system that already runs internal audits and mock inspections has the machinery for this; what it does not yet have is the drill pointed at agent-assisted decisions specifically.
And it has to survive two conditions a clean design tends to assume away: some of the workflow runs on a third-party platform whose internals the manufacturer cannot see and whose contract may end, so the record must be portable and must not depend on the vendor's retention; and the record has to be intelligible to a reader with no access to the platform at all, because that is the reader who will be holding it.
None of that is exotic. The sector already knows how to design a record backwards from the questions that will be asked of it — that is what a batch record is, and it has been doing it for longer than the observability industry has existed. What has not been done is the assembly, for this class of decision, in a form a quality organisation can operate and an investigator can read. That construction is the subject of the companion to this piece, "Decision records for agent-assisted quality systems."
The reason to do it now rather than after the guidance is the reason a deferral is a stronger argument than an exemption. The predicate rules are already binding and already general; the question is already available to an investigator today; and the guidance that eventually arrives will be written against whatever manufacturers have already built, by people reading what the industry did in the interval. That interval is open now, and it is the only period in which what you build influences what you are later measured against.