Walk into any large industrial programme in 2026 and you will find a digital twin. Walk a little further in and you will find that the word is doing an enormous amount of unearned work.

It covers a CAD model with live sensor tags. It covers a physics-accurate simulation environment used to validate control policies. It covers a dashboard someone renamed in a slide deck two years ago. These are not the same artefact, they carry wildly different obligations, and the collapse of all three into one word is how organisations end up making consequential decisions on the strength of a rendering.

The useful reframe is this. A digital twin is not a model of a thing. It is a standing claim about a thing — the assertion that, within some stated envelope, the behaviour of this representation predicts the behaviour of that physical asset. Once you say it that way, three questions follow immediately, and a twin that cannot answer them is not evidence.

The standard already says this, and almost nobody reads it that way

Before the argument, a piece of useful ammunition, because this is not a consultant's neologism. ISO 23247, the digital twin framework for manufacturing, defines a digital twin as a fit for purpose digital representation of an observable manufacturing element, with synchronization between the element and its digital representation.

Read that definition slowly, because two of its phrases are doing all the work and both are routinely skipped.

Fit for purpose. Not accurate, not high-fidelity, not photorealistic — fit for a purpose. The definition presupposes that a purpose has been stated, because fitness is meaningless without one. A twin fit for visualising throughput is not thereby fit for validating a contact-rich manipulation policy, and the standard's own language concedes as much. The field on the form that says which purpose is the field most programmes leave blank.

Synchronization between the element and its digital representation. Not a property the twin has once at build time, but a relationship that must be maintained against a physical asset that will not hold still. Synchronisation is precisely the thing that decays, and the definition makes decay a first-class failure rather than an unfortunate surprise.

So the argument in this piece is not an imposition from outside the discipline. It is what the discipline's own framework already says, read literally. The gap is between the standard's definition and what most organisations have actually built and are calling by the same name.

One word, three artefacts Same label. Entirely different claims, and entirely different obligations. Geometry + live tags A CAD model with sensor values attached to it. Claims This is the shape of the asset, and this is what its sensors read. Says nothing about how it will behave under any condition not currently occurring. Obligation Low It asserts very little, so very little can be wrong. Behavioural simulation A physics model used to predict how the asset responds. Claims Within a stated envelope, my dynamics predict its dynamics. Which means it makes assertions about states nobody has yet observed. Obligation High This is the one that ends up inside a safety case. Operational dashboard Aggregated telemetry, rendered attractively. Claims These are the numbers. The hazard Perceived authority is high because it looks like the others. Obligation Low Until someone starts deciding with it. Then see column two. ISO 23247 defines a twin as a “fit for purpose” representation, synchronised with its element. Fit for which purpose is the whole question — and it is the field most programmes leave blank. Collapsing three artefacts into one word is how a rendering ends up carrying a decision.

Question one: what exactly does it claim, and about what?

Fidelity is not a scalar, and treating it as one is the root error underneath most of what follows.

A twin can be geometrically exact and dynamically useless. It can be kinematically faithful and thermally blind. It can be excellent at steady state and wrong at every transient — which matters enormously, since transients are where most interesting failures live. Asking “how accurate is the twin?” invites a number that averages across all of these and therefore describes none of them.

The claim has a scope, and mature programmes write it down: which phenomena are modelled, which are approximated and by what method, which are absent entirely, and over what operating envelope the correspondence has actually been checked. That last clause is the one that gets skipped. A twin is frequently validated in one region of its envelope and then used across all of it, and nothing in the artefact itself records the difference.

This sounds bureaucratic until the first incident, at which point it is the only document that distinguishes two findings with entirely different consequences for everyone involved: the twin was used outside its stated envelope, or the twin was wrong. The first is an operating discipline failure with a clear owner and a clear fix. The second calls the whole programme's evidence base into question. Without a written scope, every incident defaults to the second reading, because nobody can demonstrate the first.

Question two: how far has it drifted?

This is the question almost nobody instruments, and it is where twins quietly stop being useful and start being dangerous.

Physical assets change. Bearings wear. Calibration shifts. A line is re-tooled. A fixture is replaced with a not-quite-identical part because the original had a twelve-week lead time. Someone bolts a bracket on during the third shift and it is entirely reasonable that they did, and it is nowhere in any model.

The twin does not change, unless something makes it. Every one of those divergences widens the gap between the claim and the asset — silently, incrementally, while the twin keeps rendering with completely undiminished confidence. That last property is what makes this a safety issue rather than a data-quality issue. A stale spreadsheet looks stale. A stale twin looks exactly as authoritative as a current one.

A claim that was true once The asset changes. The twin does not, unless something makes it. divergence between asset and twin year 3 year 1 year 2 commissioning stop trusting this for decisions bearing wear calibration shift line re-tooled fixture replaced bracket added, 3rd shift the twin's own confidence — unchanged throughout The fix is unglamorous and mostly missing Correspondence checks on a schedule: twin prediction vs measured reality, on the behaviours that matter, with a threshold that revokes trust when crossed. The check is cheap. Its absence is what makes the twin dangerous rather than merely stale. A twin without a drift measurement is a claim that was true once.

A twin without a drift measurement is a claim that was true once. The practice that fixes it is unglamorous and mostly missing: a defined set of correspondence checks, run on a schedule, against reality — comparing what the twin predicts to what the asset actually does on the behaviours the twin is relied upon for — with a threshold that says stop trusting this for decisions when it is exceeded.

Three design notes, because “run correspondence checks” is easy to say and easy to implement badly. Check the behaviours the claim is load-bearing for, not the ones that are convenient to measure; a twin can track throughput perfectly while its contact dynamics have drifted badly. Set the threshold before you see the data, for the same reason model validation does. And make crossing the threshold do something — an alert nobody owns is not a control, and the failure mode here is not that the check is absent but that it fires into a channel nobody reads.

Question three: what is it authorised to decide?

This is the governance question, and it is the one this argument keeps returning to, because it is the one that determines how much the other two matter.

A twin used for visualisation carries almost no obligation. Somebody looks at it; if it is wrong, somebody is briefly misinformed. A twin used to train a control policy is now upstream of a safety case: whatever the twin does not model becomes a region the policy never learned about, and the twin's fidelity limits are silently converted into the policy's blind spots. A twin used to validate that policy is inside the safety case, and its residual is the safety case's residual. And a twin wired into live operations — where its output triggers action on the physical asset — has become a control system, whatever the org chart says.

Obligations attach to the role, not the label The same artefact, used four ways, carries four completely different duties. 1 · Visualisation Someone looks at it. Obligation: almost none. Being wrong costs a misunderstanding. 2 · Training The twin generates the data a control policy learns from. Now upstream of the safety case: its fidelity limits become the policy's blind spots. 3 · Validation The twin is the environment in which the policy is tested and approved. Now inside the safety case. Its residual is the safety case's residual. 4 · Live operation Its output triggers action on the physical asset. It has become a control system, whatever the org chart says. role creep Built for rung one · quietly promoted to rung three · fidelity claim never re-examined.

The obligations attach to the role, not to the label. And the most common failure I see is role creep: a twin built for visualisation, gradually promoted into decision-making because it was there and it was good, with no one re-examining whether its fidelity claim was ever adequate for its new job.

Role creep is worth dwelling on because of how reasonable each individual step is. Nobody decides to base a safety argument on a visualisation tool. What happens is that a twin proves useful, so someone asks it a slightly harder question, and it answers plausibly, so someone asks a harder one. There is no moment where the artefact changes and therefore no moment that triggers a review. The promotion is invisible precisely because it is gradual, and the fidelity claim that was appropriate at rung one is never revisited at rung three because nobody noticed a rung had been climbed.

The control is a re-qualification trigger tied to use rather than to change: when a twin's output starts informing a new class of decision, its scope statement is re-examined against that class. This is cheap if the scope statement exists and impossible if it does not, which is the second reason to write one.

The provenance layer nobody builds until they need it

Regulated industries will eventually be asked a familiar question in unfamiliar clothes: what did the twin know, from where, and when?

Twins are assembled from CAD, sensor histories, maintenance records, material properties, vendor specifications and prior simulation runs — each with its own vintage, its own quality, and its own owner. The assembled artefact presents one uniform level of apparent confidence, which is a lie the assembly process tells by construction: nothing in the rendering distinguishes the geometry that was surveyed last month from the material property someone took off a datasheet in 2019.

What did the twin know, from where, and when? Six sources, six vintages, six owners — assembled into one confident rendering. CAD geometry vintage · quality · owner Sensor histories vintage · quality · owner Maintenance records vintage · quality · owner Material properties vintage · quality · owner Vendor specifications vintage · quality · owner Prior simulation runs under which platform version? The assembled twin One artefact. One apparent level of confidence. A decision is made using it Challenged months later, by someone adversarial Record it at assembly time Cheap, mechanical, and answerable. ISO 23247-5 calls this the digital thread. Reconstruct it afterwards Archaeology — and some of the source data was on a plant network nobody logs.

When a twin-informed decision is challenged, reconstructing that provenance after the fact is the same archaeology project that plagues agentic software — with the added indignity that some of the source data was on a plant network nobody logs, and some of it was a judgement call made verbally by an engineer who has since changed jobs.

The standards work has caught up here faster than practice has. ISO 23247-5 addresses exactly this under the name digital thread: the creation, connectivity, management and maintenance of manufacturing digital twins across the product lifecycle. The mechanism exists and is specified. What is mostly missing is the decision to use it, which is a governance choice rather than a technical one, and which is far cheaper made before the challenge than after.

The operative rule is the same one that governs agentic systems: provenance is recorded at assembly time or it is not recorded. There is no third option where you reconstruct it later at comparable cost and comparable credibility, and the organisations that will survive that conversation are the ones recording now.

The strongest objection: this is a lot of paperwork for a modelling tool

Take the objection seriously, because it is usually made by people who are building good systems and have limited time.

The argument runs: engineers have always used models, models have always been approximate, everyone competent already knows a simulation is not reality, and formalising all of this into scope statements and correspondence schedules and provenance registers is bureaucracy that will slow down the thing that is actually creating value. Most of the twins in most organisations are used sensibly by people who understand their limits. Writing it all down changes nothing except the amount of writing.

There is a lot of truth in that, and I would concede two-thirds of it. For a twin that stays at rung one, this apparatus is overhead and should be skipped. The engineering judgement of a team that built the model and uses it daily is genuinely reliable, and formal documentation is a poor substitute for it.

But the argument has a specific failure point, and it is not the competence of the team. It is the durability of that competence. The knowledge that keeps an informally governed twin safe lives in a small number of heads: which parts of the model to trust, which numbers were always rough, what changed in 2024 and why. That knowledge does not survive a reorganisation, a retirement, an acquisition, or the promotion of the twin to a role its original authors never contemplated. The documentation is not for the people who built it. It is for the people who inherit it, and for the twin's third year, which is precisely when the drift the diagram shows has accumulated and the original team has moved on.

So the honest scoping is: rung one, don't bother. Rung two and above, and especially any twin whose output reaches a safety argument or a live actuator, the apparatus pays for itself the first time anyone asks a hard question — and the whole point is that you do not get to choose when that is.

A fidelity profile, not a fidelity score

If fidelity is not a scalar, the practical question is what replaces the scalar. The answer is a profile: a short set of dimensions, each rated separately, because they fail independently and are relied upon independently.

Geometric. Does the representation match the physical shape, placement and tolerances? This is the dimension twins are usually best at, because it comes from CAD, and it is also the one whose accuracy is most often used as a proxy for the others. A geometrically immaculate twin tells you almost nothing about whether its dynamics transfer.

Kinematic. Do the motions, reachability and joint limits match? Usually good, usually derived from vendor data, and usually valid until someone changes an end effector or a fixture and the twin keeps the old one.

Dynamic. Do forces, inertias, accelerations and settling behaviour match? This is where the fidelity claim starts being genuinely expensive to substantiate, and where the difference between steady-state accuracy and transient accuracy usually appears.

Contact. Do friction, compliance and deformation match? Almost always the weakest dimension, for the reasons the reality-gap literature sets out at length, and almost always the one a manipulation safety case actually depends on.

Thermal and wear. Are effects that develop over minutes and months represented at all? Frequently absent entirely — which is defensible, and becomes indefensible only when the twin is used to reason about a duty cycle or an asset's third year without anyone noting the omission.

Temporal. Do latencies, cycle times and synchronisation behaviour match? This dimension is distinctive because a twin can be perfect on every other axis and still mislead: a control interaction that is stable at simulated timing can oscillate at real timing.

Rating six dimensions instead of quoting one number takes an afternoon and changes the conversation permanently, because it forces the sentence nobody wants to write: this twin is strong geometrically and kinematically, adequate dynamically, and does not represent contact or thermal behaviour at all. That sentence is what a reviewer needs. It is also, usually, what the engineers already believe and have never been asked to state.

What a scope statement looks like on paper

Concretely, for one twin, on one page.

Name the asset and the twin's declared purpose, in the language of decisions rather than of technology: this twin exists to validate cycle-time and collision-free path claims for cell 4, and to visualise throughput. Then the fidelity profile above, six lines. Then the envelope: the ranges of speed, payload, temperature and configuration over which correspondence has actually been checked — with the emphasis on actually, since the checked envelope is routinely narrower than the used envelope and the difference is invisible without this line.

Then the exclusions, stated positively rather than by omission: this twin does not represent contact compliance, tool wear, or operator presence. Then the correspondence regime — which behaviours are checked, how often, by what measurement, against what threshold, and who is notified when the threshold is crossed. Then the role register: what decisions this twin's output currently informs, updated when that list changes, which is the artefact that makes role creep visible.

Finally, provenance: the sources the twin was assembled from with their vintages, and the platform and version it runs on. One page. The programmes that have it treat it as unremarkable; the programmes that do not tend to discover it is the document everyone wanted the week they cannot produce it.

Who owns the claim?

Every control described here presumes an owner, and the ownership question is where most of these programmes actually break — not on technique, on org design.

A twin typically has three de facto custodians and no accountable one. Engineering built it and understands its physics. Operations uses it daily and knows where it lies. IT hosts it, patches it, and has no view on whether its contact model is adequate. Risk has usually never heard of it, because it arrived as a tool rather than as a system, and tools do not go through the intake process that systems do. The result is that when the twin's correspondence threshold is crossed, or its role creeps, or a platform version changes, the question of who decides lands in the gap between three functions.

The ownership that matters is ownership of the claim, which is a different thing from ownership of the software. The person who owns a twin's claim is the person who is prepared to say: within this envelope, this representation predicts this asset, and I will be asked about it if it does not. That is a named individual with the standing to stop a deployment, not a team inbox.

Two practical consequences. First, the claim owner should be the person who can actually revoke trust — which usually means engineering or operations rather than IT, because the decision is about physics rather than about availability. Second, the moment a twin's role climbs the ladder, the claim owner should change or be re-confirmed deliberately, because the person qualified to stand behind a visualisation claim is not necessarily the person qualified to stand behind a validation claim.

The diagnostic question is simple and slightly rude, which is why it works: if this twin were wrong tomorrow in a way that mattered, whose judgement was it? If the room offers a function rather than a person, the claim is unowned, and every control above is decorative.

The model-risk analogy, and its limits

Everything above will feel familiar to anyone who has worked in model risk, and the parallel is worth drawing explicitly because it imports a decade of hard-won practice — carefully, because it also imports assumptions that do not hold.

What transfers cleanly: the idea that a model has a documented purpose and a validated range of use; that use outside that range is a control failure with an owner; that models are inventoried, so somebody can answer how many there are and what depends on them; and that a change to the model or its inputs triggers re-examination rather than being absorbed silently.

One transfer needs care. US supervisory guidance on model risk was substantially revised in April 2026 — OCC Bulletin 2026-13 and the parallel Federal Reserve and FDIC issuances — which supersedes the older framework practitioners still cite reflexively, states that generative and agentic AI models are not within its scope, and does not set out enforceable standards. Read as an exemption, that is a mistake in both directions: it does not license financial institutions to skip governance on AI models, and it certainly says nothing about industrial twins, which were never in scope of banking guidance to begin with. What it offers is a vocabulary and a set of proven controls, not an authority to point at.

And one thing does not transfer at all. A financial model that is wrong produces a bad number, which someone downstream may catch. A twin that is wrong about contact dynamics, sitting at rung three or four, produces a machine that moves. The remediation window that makes detect-and-correct workable for model risk is exactly what physical deployment removes.

Whose formats is your process knowledge in?

A last consequence, easy to miss because it is commercial rather than technical.

The industrial world has spent decades building genuine sovereignty over its physical processes — tooling, methods, and the accumulated understanding of what actually goes wrong on this line — encoded in people, procedures and drawings. As validation moves into simulation, that process knowledge moves with it: the twins, the scenarios, the tuned parameters, the hard-won library of failure cases. That asset is increasingly expressed in a platform vendor's formats, and improves that vendor's ecosystem.

This is not an argument against the platforms, which are excellent and whose alternatives are years behind. It is an argument for noticing what is happening and pricing it. For a national industrial programme — India's manufacturing build-out, the Gulf's giga-projects and automated ports — this is the same fork the AI-model conversation reached last year, arriving in the physical economy: own the layer where your process knowledge lives, or rent it.

The practical mitigation is modest and worth doing even if you never exercise it: keep twins, scenarios and validation criteria in formats you could port, and know what it would cost to move. Not because you plan to, but because a dependency whose exit cost you have never estimated is not a dependency you have actually accepted — you have only defaulted into it.

A short test you can run this week

Take your programme's most consequential twin and ask four questions. What does it claim, and over what envelope? When was that claim last checked against the physical asset, and by what measurement? What decisions does its output currently influence — including the informal ones, which is where role creep hides? And if a decision it informed were challenged tomorrow, what could you actually show?

Four questions, one meeting. As with any question set worth asking, the pattern of the silences is the finding. A programme that answers one and two crisply and stumbles on three has a role-creep problem. A programme that answers one through three and cannot answer four has a provenance problem it can start fixing today at low cost. A programme that stumbles on all four is operating an asset it does not understand, which is worth knowing before someone else discovers it.

Where this argument is weakest

Two places worth naming.

First, “measure drift” is straightforward for a twin of a single machine with instrumented behaviour and genuinely hard for a twin of a whole line or facility, where the correspondence checks multiply combinatorially and the thing you would want to compare against is itself only partially observable. I have described the discipline; I have not solved the measurement problem for large-scale twins, and I am not aware of anyone who has.

Second, the argument for writing scope statements rests on a claim about how incidents get adjudicated that is grounded in adjacent domains — model risk, functional safety — rather than in a body of settled physical-AI cases, of which there are few. It is a well-supported extrapolation rather than an observed regularity, and it should be read as such.

None of this is an argument against digital twins. The opposite: twins are the most powerful instrument the physical industries have acquired in a generation, and the programmes getting them right are compounding advantages competitors will take years to match. But the value is in the claim, the claim decays, and a claim nobody maintains is worse than no claim at all — because people believe it.