The companion teardown to this piece, The inspection arrives years after the agent has gone, argues that agent-assisted quality workflows are leaving behind spans, tool calls and model outputs in a sector where the question arrives four years later and asks what a decision rested on. I am going to take that as established. What follows is the object I would build against it, and — because a primitive that only lists its own virtues is worthless — a considerably longer account of what it does not do.
Start with the decision, because the design falls out of the question rather than out of the architecture.
A deviation is raised on a packaging line and an agent-assisted workflow classifies it as minor. Four years later an investigator asks which version of the procedure was in front of the person who signed, and on what basis it was classified minor. Work backwards from that sentence and the object specifies itself. The investigator needs: what was decided; what the decision consumed, in the versions that were actually in force at the time; why the conclusion follows from that; what else was on the table and why it was rejected; and who accepted it, under what authority, having seen what. Five things. Nothing about that list is novel to this sector — it is the shape of an investigation report, an impact assessment, a change control record. The novelty is only that it now has to be produced for a class of decision where part of the work was done by a machine that does not naturally produce any of it.
The packaging-line deviation, the classification and the four-year inspection are a constructed illustration carried forward from the companion piece, assembled from patterns documented in published regulatory guidance, industry practice literature and vendor deployment documentation. It is not a report of any real site, product, batch, deviation, manufacturer or deployment, and no part of this piece describes client work.
The ground the design stands on
Four premises come across from the teardown. I am not going to re-argue them; each is defended there, and if any of them is wrong the design below is wrong with it.
- The record is constructed at the decision, not derived from telemetry afterwards. Derivation is reconstruction, and reconstruction is the failure mode. If the record can only be assembled by an engineer querying a trace store, then it does not exist once the trace store rotates — and it is unavailable to the quality organisation in the meantime.
- Effective versions are bound at retrieval. The identifier of a document is not the same as the state of that document at a moment. Resolving "what did procedure QA-114 say in March of the year before last" from a document management system's history afterwards is possible, sometimes, and it is an inference. Capturing the effective version when the document is consumed is a fact.
- The record has to be readable by someone with no access to the platform. The reader who matters is holding the document in a room, possibly years after the platform contract ended. A record whose meaning depends on the vendor's console is not a record; it is a query.
- Retention runs on the product's clock. The batch-associated records are kept past the product's expiration date under the predicate rules. The decision record is a batch-associated record in substance, whatever the platform's storage tier thinks it is, and it should inherit that floor rather than the observability system's.
A fifth premise is specific to agent assistance, and it is the one that generates most of the design's awkwardness. The approver's contribution has to be recorded for what it actually was. In a conventional investigation, the person who signs the report is usually the person who did the investigating, or is signing off work whose author is separately named — and either way, the signature attests to a decision they participated in. In an agent-assisted workflow, the person who signs may have authored the reasoning, may have re-derived it independently and agreed, or may have read a fluent draft and accepted it. Those three are different facts about how much human judgement is behind the disposition. Today they produce identical signatures, and a record that does not distinguish them is asserting something it does not know.
The object: five sections, and the one that carries the weight
1 — The decision, in the quality system's vocabulary. Not "model output: minor". The decision as the quality management system understands it, phrased as the procedure phrases it, attached to the record it modifies, with the moment it was made. This sounds trivial and is the section most often got wrong, because platform-native records describe the platform's action rather than the regulated act. A record that says a classification step returned a label is describing an inference; a record that says a deviation was classified minor is describing a decision. Only the second one has a place in the quality system.
2 — Inputs consumed, each with the version in force when it was consumed. Every document, specification, dataset, prior deviation and equipment record that entered the decision, each carrying an effective-version identifier captured at the moment of retrieval. This is the section that requires engineering work rather than discipline: the retrieval layer has to ask the document store for a version and record the answer, and most retrieval layers were built to return relevant text quickly and were never asked to. Note what is being captured and what is not — this is what came back, not what the model attended to. The second is not available, and I will come to that.
3 — Basis: the reasoning connecting those inputs to that decision. Written so that a reader with no access to the platform can follow it, which in practice means written in the terms of the procedure rather than the terms of the model. Where an agent drafted the basis, that fact belongs in section five rather than being hidden here. The test for this section is not whether it is persuasive; it is whether a competent reader could disagree with it. A basis that cannot be disagreed with is a restatement of the conclusion.
4 — Alternatives considered and rejected. This is the load-bearing section. What else was available — the other classifications, the other candidate root causes, the other dispositions — and why each was not taken. It is load-bearing for a reason that has nothing to do with regulation and everything to do with epistemics: a decision that never had alternatives was not a decision, it was an output. This is the section that lets an investigator test the judgement rather than merely observe it, and it is the section that most obviously distinguishes a workflow where a human exercised judgement from one where a human agreed with a machine. It is also, for exactly that reason, the section most at risk of degrading into boilerplate, which is the second failure mode in the bill below.
5 — Approver: a named person, the authority, and what they were shown. The natural person, the authority they were acting under in the quality unit's mandate, and — this is the part that is new — the state of the artefact at the moment they approved it. If the reviewer approved a draft that the agent produced, the record says so and captures the draft as approved. If the reviewer materially altered it first, the record captures both states. If the reviewer re-derived the conclusion independently, the record says that instead. None of these is disqualifying; the point is that they are different, and the difference is the thing an inspection will care about most and the thing today's signature obscures completely.
Five sections, and the completeness rule is that all five are present or the decision is not recorded. That is deliberately harsh, and it is what makes this a control rather than a template. A record with four sections filled and one blank is worse than no record at all, because it looks like evidence.
The one thing this design cannot do
Before the mechanics, the honest limitation, stated first rather than buried at the end where limitations usually go.
This design does not record the model's reasoning, and it cannot. The mechanism that produces a language model's output does not decompose into a basis that can be extracted, and an explanation generated afterwards — by the same model, asked why it said what it said — is a second output rather than a record of the first. Treating a generated rationale as evidence of the actual basis would be introducing a fabrication into a regulated record, which is a considerably worse failure than the one this piece is trying to fix. Any vendor offering "explainability for your audit trail" on that basis should be asked, precisely, whether the explanation is a causal account of the computation or a plausible narrative generated after it. The honest answer is the second one.
So what does the record contain in place of the model's reasoning? Two things, both weaker and both true. First, the decision procedure's inputs: what was retrieved, in what versions, and what the model output was — which supports a competent reader in judging whether the conclusion is defensible on the material, without claiming to describe how the conclusion was reached. Second, the human reviewer's own basis, stated by the reviewer. That is a real basis belonging to a real person who can be asked about it, and it is the thing that actually carries the decision in a regime that terminates in accountable people.
This makes the reviewer's job larger, not smaller, and the design should not pretend otherwise. If the reviewer must state a basis in their own terms rather than signing a draft, then some of the time saved by the drafting is spent again. That is a real cost, it lands on the exact people the automation was sold to relieve, and it is the sharpest practical objection to this entire design. My answer is not that the cost is small. It is that the cost is the thing being bought: a signature on a fluent draft is cheap because it is not evidence of anything, and the decisions where this matters are a minority of the volume. The design should therefore be scoped by decision class rather than applied uniformly, which is the first item in the build order below.
The data model
The types below are the reference shape, written to be read rather than dropped into a validated environment. Two things in them are doing real work: the version binding is a required field on every consumed input, so a retrieval that did not capture a version cannot produce a well-formed record; and the approver's contribution is a discriminated union rather than a boolean, so the three cases cannot collapse into one.
The decision record, as a type
Three files. The first is the record itself, with the completeness rule expressed in the type system rather than in a procedure. The second is the retrieval-time version binding — the piece of engineering that has to exist upstream or the record cannot be built. The third is the drill: the query an investigator would run, written against the record, so that a failure to answer is a compile-or-runtime failure rather than an awkward meeting.
Note that ApproverContribution is a union of three cases rather than a flag. The three are genuinely different facts about how much human judgement is behind the disposition, and a record that cannot distinguish them is making a claim it has not earned.
/** A document, specification or dataset as it stood when it was consumed. */
export interface ConsumedInput {
/** Stable identity of the thing — e.g. a procedure number or a dataset key. */
readonly documentId: string;
readonly title: string;
/** The version in force at the moment of retrieval. Required: a record whose inputs
* carry no version is not a decision record, it is a bibliography. */
readonly effectiveVersion: string;
/** When the version was resolved. Distinct from the decision time. */
readonly retrievedIso: string;
/** How it entered the decision: retrieved by the agent, attached by the reviewer,
* or already bound to the record under investigation. */
readonly channel: "agent-retrieval" | "reviewer-attached" | "record-bound";
}
/** What was rejected, and why. Empty is not allowed — see requireAlternatives below. */
export interface RejectedAlternative {
readonly option: string;
readonly reasonRejected: string;
}
/** The three cases. They are not interchangeable and the type refuses to let them be. */
export type ApproverContribution =
| { readonly kind: "authored"; readonly note: string }
| {
readonly kind: "re-derived";
/** The reviewer reached the conclusion independently before seeing the draft. */
readonly note: string;
}
| {
readonly kind: "accepted-draft";
/** The artefact exactly as presented, so that what was approved is recoverable. */
readonly draftAsPresented: string;
/** Any material alteration the reviewer made before approving. */
readonly alterations: readonly string[];
};
export interface Approver {
readonly directoryId: string;
readonly displayName: string;
/** The mandate under which they may approve this class of decision. */
readonly authority: string;
readonly approvedIso: string;
readonly contribution: ApproverContribution;
}
export interface DecisionRecord {
readonly recordId: string;
/** The record in the quality system this decision attaches to. */
readonly attachesTo: { readonly system: string; readonly recordId: string };
/** 1 — the decision, phrased as the procedure phrases it. */
readonly decision: string;
readonly decidedIso: string;
/** 2 — everything the decision consumed, versioned. */
readonly inputs: readonly ConsumedInput[];
/** 3 — the basis, in terms a reader without the platform can follow. */
readonly basis: string;
/** 4 — the load-bearing section. */
readonly alternatives: readonly RejectedAlternative[];
/** 5 — who accepted it, and what their contribution actually was. */
readonly approver: Approver;
/** An optional pointer into the telemetry. Useful while it lasts; never required. */
readonly runReference?: { readonly correlationId: string; readonly platform: string };
/** Inherited from the batch, not from the log. */
readonly retainUntilIso: string;
}
/**
* Completeness fails closed. A record missing any section is not a weaker record —
* it is a record that will be read as evidence and is not.
*/
export function isComplete(record: DecisionRecord): boolean {
return (
record.decision.trim().length > 0 &&
record.inputs.length > 0 &&
record.inputs.every((i) => i.effectiveVersion.trim().length > 0) &&
record.basis.trim().length > 0 &&
record.alternatives.length > 0 &&
record.approver.authority.trim().length > 0
);
}These files are a reference shape for reading, not a validated artefact. Anything deployed into a GxP environment is subject to the manufacturer's own computerised system validation and change control, and nothing in an essay discharges that.
Holding it apart from the telemetry
The second structural decision is where the record lives, and it is the one most likely to be got wrong by an engineering-led implementation, because the natural instinct is to enrich the trace.
Two planes with two clocks. The telemetry plane keeps its engineering retention and its engineering purpose: spans, tool call records, model outputs, cost and latency metrics, retained for as long as debugging and cost control justify. Nothing here argues for retaining telemetry longer, and estates that respond to this argument by extending trace retention to seven years have bought storage instead of evidence.
The record plane belongs to the quality organisation, holds the decision records, and inherits the retention floor attached to the batch. The link between them is a single optional reference — a correlation identifier that lets an engineer find the run while the run still exists. It is a pointer, not a dependency. If the trace store rotates, if the platform is replaced, if the vendor relationship ends, the record is unaffected, because everything load-bearing was copied into it at the moment of the decision rather than referenced out of it.
That portability is not a nicety in this sector. Platform contracts are renegotiated on cycles far shorter than product lifetimes, and the exit is exactly when evidence goes missing — not through bad faith, but because a terminated contract terminates access, and the data export nobody scoped at signing turns out to be a JSON dump of spans. A decision record that can be exported as a document, read by a person and filed alongside the batch record has no exit problem, because it was never inside the platform in the sense that matters.
The drill: making the record a control rather than a document
Every property above can be satisfied on paper and fail in practice. The record specification can be right and the records can be filled with boilerplate. The version binding can be implemented and quietly bypassed by a retrieval path someone added later. The only way to find out is to ask the inspection question of a real past decision, on a schedule, and write down what could and could not be answered.
The mechanics are deliberately unremarkable, because a quality organisation already runs internal audits and mock inspections and the point is to use that machinery rather than build a parallel one.
- Sample from the live population, weighted toward the old. A drill run against last month's decisions tests the platform, because the platform's data is still there. The frame has to include decisions old enough that the telemetry has already rotated and, ideally, old enough that at least one person involved has moved on. That is the condition the record was built for, and it is the only condition under which the test is real.
- Pose the question as an investigator would. Not "can we retrieve the trace" but "which effective versions informed this, on what basis, approved by whom under what authority". The wording matters because the engineering-framed version of the question is answerable by a system that fails the real one.
- Score into three outcomes, not two. Answerable from the record is the pass. Not answerable is a defect in the record specification. The middle outcome — answerable only by asking someone who remembers — is the important one, and organisations that score pass or fail will silently record it as a pass. It is a latent failure, because memory is the channel that expires next.
- Route each outcome to a different remediation. Pass: nothing. Memory-only: identify the field that was missing and capture it at the decision from now on. Not answerable: change the record specification, which is a change-controlled act and should look like one.
- Close the loop. Re-drill the same class after the remediation. A periodic exercise with no feedback edge is a report; what makes this a control is that the next drill tests whether the previous remediation worked.
The objections, at full strength
Three objections deserve to be stated properly, and the first is the one I would make.
Computerised system validation and the data-integrity regime already govern this, so you are proposing a second rulebook. Fully conceded on the first half and answered on the second. Validation and data integrity are real, mature and better developed here than anywhere else in the private economy, and nothing above competes with them. The distinction the companion piece draws is that validation establishes a system performs as intended across its intended use, and data-integrity expectations set quality attributes on the records a system creates — neither specifies which records must exist for a class of decision that did not exist when the procedures were written. The decision record is not a second rulebook. It is a record type, added to a quality system that already has dozens of record types, and it should be introduced through the same change control as any other. If your validation package already requires this record and your procedures already specify its content, you are done and this piece is describing your estate rather than proposing something to it.
This is documentation theatre with a compliance cost and no product benefit. A serious objection, and it lands hardest against a uniform rollout. The answer is scoping rather than volume: the record is worth building for decisions where the outcome determines the depth of an investigation, the disposition of product, the reportability of an event or the content of a submission, and it is not worth building for the long tail of routine categorisations that no investigator will ever pull. A design that applies to every decision will be complied with formally and hollowed out substantively within two quarters, which is worse than not doing it. The product benefit, where there is one, is population scoping: the estate that can say which decisions consumed a superseded procedure can bound a defect in hours instead of re-reviewing a year of them.
The reviewer bottleneck: you have re-imposed the work the automation removed. Partly true, and I have already conceded the mechanism above. Asking a reviewer to state a basis in their own terms costs some of the time the drafting saved. Three things bound the cost. It falls only on the scoped decision classes. It is bounded by the fact that a reviewer who genuinely cannot state a basis has discovered something important about that decision rather than about the process. And the alternative — a signature on a fluent draft, indistinguishable from a signature on an independently derived conclusion — is not a cheaper version of the same evidence, it is the absence of evidence at a lower price. But I have no measurement of the added review time, and anyone quoting one to you should be asked for their methodology.
The bill: seven ways this design fails
A primitive that lists only its properties is marketing. These are the failure modes I would expect, in roughly the order I would expect to meet them.
- The alternatives section degrades into boilerplate. This is the most likely failure and the most damaging, because the section is load-bearing. Under volume, "major — not supported by the reconciliation data" becomes a phrase pasted into every record, and the record now asserts that alternatives were considered when they were not. The drill catches it only if the drill reads the content rather than checking the field is non-empty, which means the scoring has to be done by someone competent in the subject matter.
- The basis becomes a restatement of the conclusion. Closely related and harder to detect automatically. "Classified minor because the deviation is minor in nature" satisfies every structural check. The only test I know is whether a competent reader could disagree with the stated basis, and that test cannot be automated.
- Version binding is bypassed by a path added later. The retrieval layer is fixed, the records are good for a year, and then someone adds a second retrieval path for a new document class under schedule pressure. Nothing fails visibly; the records simply start carrying inputs whose versions were resolved afterwards, which is a reconstruction wearing a record's clothes. This wants a technical control at the record-construction boundary rather than a procedural one.
- Scoping drifts. Decision classes in scope are chosen at design time and the world moves. New workflows arrive, an existing workflow's outputs start feeding a reportability judgement, and nobody revisits the scope because scope review is nobody's task. The scope list should be a change-controlled artefact with an owner and a review trigger, not a table in an implementation document.
- The approver contribution field is filled in by the platform rather than the person. If the workflow defaults to "accepted-draft" and the reviewer never selects otherwise, the field records the platform's assumption rather than the reviewer's act. If it defaults to "re-derived", it is worse — it is a fabrication in a regulated record. The field must be an explicit act by the person, and if that is unacceptable to the workflow designers then the honest choice is to record only "accepted-draft" and stop claiming more.
- The record is designed against today's inspection questions. I have specified five sections from the questions I understand investigators to ask. Inspection practice changes, and it will change faster now than it has for a decade as investigators develop expectations about AI-assisted work. A record designed against the current question set may be missing the section that matters in 2029, and the drill will not find that, because the drill asks the questions the designers thought of. The mitigation is to keep the record extensible and to update the drill's question set from whatever the sector's inspection findings start to show — which is a dependency on information that does not exist yet.
- The record becomes the thing that is audited instead of the decision. The endpoint of every documentation control is that people optimise the artefact. Once the drill is a metric, the incentive is to produce drill-passing records rather than good decisions, and the two come apart in exactly the cases where judgement was hardest. I do not have a structural answer to this beyond the usual one — that the drill should be run by people whose job is the product rather than the paperwork — and I am not confident it is sufficient.
The costs are lower than the failure modes but should be stated. Engineering: the version-binding change at retrieval, a record store the quality organisation controls, and a construction boundary that fails closed. Validation: the record type and its workflow go through the manufacturer's normal computerised system validation and change control, which is not a small line item and is not avoidable. Storage: negligible, because a decision record is a few kilobytes of structured text and the thing being retained for years is the record rather than the traces. Review time: increased on the scoped classes by an amount I cannot quantify and have not seen quantified anywhere. Every one of those figures that is unmeasured is marked unmeasured deliberately.
What would falsify this
- A measurement showing that reviewer recall carries the reconstruction adequately. If it turns out that reviewers of agent-drafted work remember the basis as well as authors of their own investigations, at inspection-relevant intervals, then the interview mechanism still works and this design is solving a problem that decays more slowly than I claim. I know of no such measurement in either direction.
- An instrument that specifies a different record. If FDA issues guidance facing AI in GxP operations that names a different structure, the specified structure wins and this design becomes an interesting parallel. That is a good outcome and the reason to build now is that it is the interval in which what manufacturers build shapes what gets specified.
- Evidence that the alternatives section cannot be kept honest at volume. The first failure mode is the design's weakest point. A study — or an internal drill programme with published methodology — showing that alternatives sections become boilerplate in a large majority of records within a year would mean the load-bearing section does not bear load, and the design would need a different discriminator between judgement and output.
- A platform that genuinely produces a causal account of the model's contribution. I have argued this is not available and that generated rationales are a second output. If that changes — if some mechanism produces an account of the computation that can be verified rather than merely read — then the second section of the record could carry considerably more, and the design should be revised toward it rather than defended.
A build order, and what it deliberately omits
If I had one quarter and an existing agent-assisted workflow in a quality system, this is the sequence, and the omissions are as deliberate as the inclusions.
- Scope the decision classes. One workshop with the quality unit, producing a list of decision types where the outcome determines investigation depth, product disposition, reportability or submission content. Everything else is out of scope and stays out until someone argues it in. This step is first because doing it last is how uniform rollouts happen.
- Fix retrieval so it binds versions, and make unversioned retrieval fail. This is the only genuinely non-trivial engineering change and everything downstream depends on it. It should fail closed from the first day, because a permissive period produces a corpus of records that look complete and are not.
- Add the record type to the quality system through normal change control. Not a new system. A record type, in the system the organisation already validates, audits and inspects against. The moment this becomes a separate platform, it acquires a separate retention policy and a separate exit problem and the design has defeated itself.
- Change the approval step so the contribution is an explicit act. Three options, chosen by the person, with the draft-as-presented captured in the third case. This is a small interface change with a disproportionate effect on what the record can honestly claim.
- Run the first drill against decisions from before the programme started. They will fail. That is the baseline, and having it in writing is what makes the later drills mean something.
What that order omits: any attempt to instrument the model, any explainability layer, any change to the telemetry, any new storage platform, and any claim about the model's reasoning. Those omissions are the design. What is left is a record type, a retrieval fix, an approval change and a drill — four things, none of them novel to a sector that has been designing records backwards from inspections for longer than the observability industry has existed.
The companion teardown ends by saying the interval is open. This is what I would put into it.