Every physical-AI programme now runs on the same premise: train and validate in simulation, deploy to hardware. The premise is sound, the tooling is genuinely extraordinary, and the economics are unarguable. You cannot crash ten thousand real forklifts to learn a policy, and nobody should want to.

But there is a sentence in the 2026 reality-gap literature that ought to be read aloud in every safety committee that has ever signed off a simulated validation. A safety claim supported in simulation may be weakened or invalidated when the same policy is deployed on hardware with different sensing noise, contact physics, actuator limits or object variation.

Read that as an engineer and it is a familiar caveat — the kind of thing you nod at and move past, because of course the simulator is not the world. Read it as a risk officer and it is something else entirely: the evidence you approved and the system you deployed are not the same object.

That is the whole argument of this piece, and everything below is an attempt to take it seriously rather than to state it dramatically. The engineering here is good. The people doing it are not careless. The problem is that the field has built a magnificent apparatus for making the gap smaller and almost nothing for saying how big it still is — and those are different deliverables, wanted by different people, for different reasons.

The gap is a stack, not a number

The first thing to get right is that “the reality gap” is a singular noun covering at least five distinct phenomena, with different physics and different failure modes. Treating it as one quantity is how programmes end up with a single fidelity score that means nothing in particular.

Actuation. Real motor drivers introduce microsecond delays, current caps, and dead zones near zero velocity. Each discrepancy is individually small and they compound precisely during the fast, precise movements where safety margins are thinnest. The gap is not uniformly distributed across the operating envelope; it concentrates exactly where you would least like it to.

Sensing. Noise, latency, dropout and calibration drift, none of which any simulator models exactly. The important property here is temporal: sensing degrades differently as hardware ages, which means the sensing component of the gap is not a constant to be measured once at commissioning.

Contact. Friction, compliance, deformation. This is the cruellest item on the list, because the physics that is hardest to simulate faithfully is also the physics that manipulation depends on most. A policy that is robust in free space can be brittle at exactly the moment it touches something.

Variation. Real objects are not the objects in the asset library, real floors are not level, real lighting is not the render. And this one has a multiplier that the others do not: a policy validated once is typically deployed many times, and every site contributes variation the validating twin did not contain.

Co-presence. The human who walks into the cell, whose behaviour was never in the training distribution at all. Not under-represented — absent. This is the item that turns a performance question into a safety question.

The gap is a stack, not a number Five discrepancies, each with its own physics and its own failure mode. 1 · Actuation Microsecond driver delays, current caps, dead zones near zero velocity. Compounds during the fast, precise motion where margins are thinnest. 2 · Sensing Noise, latency, dropout, calibration drift — modelled by no simulator exactly. Degrades differently as the hardware ages. The gap is not static. 3 · Contact Friction, compliance, deformation. The physics hardest to simulate is the physics manipulation depends on. 4 · Variation Real objects are not the asset library. Real floors are not level. A policy validated once inherits every site it is later deployed to. 5 · Co-presence The human who walks into the cell. Behaviour that was never in the distribution at all. Domain randomisation and higher fidelity narrow every one of these. Neither of them tells you how much is left.

The comprehensive 2026 survey of the field — “The Reality Gap in Robotics: Challenges, Solutions, and Best Practices,” from researchers at the University of Zurich, NVIDIA and the University of Washington — frames the discipline's response as a systematic pipeline: high-fidelity physics, domain randomisation, synthetic data generation, real-to-sim transfer, sim-real co-training, hardware-in-the-loop validation, careful edge deployment. That pipeline is good engineering. It demonstrably narrows the gap across locomotion, navigation and manipulation.

It does not close the gap, which everyone in the field says openly. The part that matters for governance is subtler and much less discussed: it does not tell you how much gap remains.

Why this is an evidence problem, not a fidelity problem

Here is the trap that regulated deployers keep walking into, and it is a trap precisely because the thing that springs it looks so much like the thing you wanted.

Simulation produces something with every surface property of evidence. It is quantitative. It is reproducible — arguably more reproducible than any physical test campaign, since you can re-run the identical scenario indefinitely. It is auditable, in the sense that the artefacts exist and can be inspected. It is generated at a scale no physical programme could approach: ten million simulated hours, a 99.97 percent success rate, a full sweep of edge cases including ones you would never dare to stage with real hardware and real people.

It has every property of proof except the one that counts. It is proof about a model of the machine, and the delta between that model and the machine is precisely the quantity nobody measured.

It has every property of proof except one Simulated validation is not weak evidence. It is strong evidence about the wrong object. What the simulation gives you Quantitative Reproducible Auditable Scale no physical campaign can match Edge cases you would never dare stage 10,000,000 simulated hours 99.97% success full distribution of modelled edge cases All true. All about a model of the machine. What the safety case needs Quantitative Reproducible Auditable Scale Coverage of the modelled distribution A stated bound on the residual How far may the real system deviate from the model that was actually tested? Nobody measured this. It is the whole question. One row separates the two columns. A simulated result plus a bounded, stated gap is evidence. A simulated result alone is a demo with excellent production values.

Compare this against what a functional-safety regime actually expects. Frameworks like ISO 13849 and IEC 61508 are built on characterising a system's failure behaviour and demonstrating that a required performance level holds. Underneath sits an assumption so basic it is rarely stated: that the system's behaviour can be characterised in the first place — enumerated, bounded, reasoned about component by component.

A learned policy interacting with contact physics does not offer that assumption for free. Its behaviour was characterised statistically, over a distribution, in an environment that is itself an approximation. This is not a reason to refuse learned policies in physical systems; it is a reason to be precise about what a simulated result licenses you to claim.

So the honest question for any simulated validation is not “what did the simulation show?” It is three harder ones. How much of the real system's behaviour does the simulation actually cover? How large is the residual discrepancy on the behaviours the safety case depends on? And what happens on the hardware when that residual is exceeded?

The steelman: simulation is often better than the physical test it replaced

An argument like this one is easy to read as anti-simulation, and it would be wrong to leave that impression standing, so let me put the strongest version of the other case.

Physical validation campaigns have their own severe coverage problem, and it is usually worse. A physical campaign runs the scenarios someone thought of, on the hardware available, in the conditions that obtained during the test window, at a sample size constrained by cost and calendar. It systematically under-tests rare and dangerous cases, because rare cases are rare and dangerous cases are dangerous. Its reproducibility is poor. Its edge-case coverage is frequently a handful of runs. Nobody writes a residual for it either — the implicit claim, that the tested hardware represents the fleet and the test conditions represent the field, is every bit as unquantified as a simulator's fidelity assumption, and is usually less examined because it is less visible.

Against that baseline, a physically accurate twin is not a downgrade in rigour. It is very often a large upgrade, and a programme that replaces a thin physical campaign with a broad simulated one has probably made its system safer, not less safe.

Concede that fully. It sharpens rather than blunts the argument. The problem is not that simulated evidence is weaker than physical evidence — frequently it is stronger. The problem is that the move to simulation was justified on coverage grounds, and coverage of the modelled distribution was then quietly allowed to stand in for coverage of reality. The residual question was not newly created by simulation. It was newly made answerable by simulation, and then not answered.

The research that is actually addressing this

A line of 2025–26 work is trying to quantify the gap rather than merely narrow it, and I think it is the most under-noticed development in the field for exactly the reason that it is unglamorous: it does not make robots do anything new.

The first strand is neural simulation-gap functions (Sangeerth and Jagtap, 2025), which learn a function that formally measures the discrepancy between a tractable mathematical model and a high-fidelity simulator. The paper's own framing is the useful part: with such a function in hand, existing model-based tools can be used to design a controller for the mathematical system while formally guaranteeing a decent transition from simulation to the real world. The guarantee is claimed across the state space, including at points never seen in training — which is the property that distinguishes a bound from a benchmark.

The second strand extends this to stochastic settings (Sangeerth, Lavaei and Jagtap, 2026): stochastic simulation-gap functions that formally quantify the gap between an approximate mathematical model and a high-fidelity stochastic simulator, supporting a controller design that satisfies the desired specification in the simulator with high confidence.

Narrowing the gap ≠ bounding the gap Two research tracks. Only one of them produces evidence. Narrow it High-fidelity physics Domain randomisation Synthetic data generation Real-to-sim transfer Sim-real co-training Hardware-in-the-loop validation Mature · widely deployed · genuinely effective This is good engineering, and it works. Output a smaller gap of unknown size Bound it Neural simulation-gap functions · 2025 Learns a formally guaranteed bound on the discrepancy across the state space, including at unseen points. Stochastic simulation-gap functions · 2026 Carries a probabilistic guarantee from the designed controller to the high-fidelity simulator, with high confidence. Early · demonstrated on small systems (a Mecanum bot, a pendulum) Output a gap of stated size Only the right-hand track produces an artefact a safety case can consume. It is also the least-discussed work in the field, and the earliest.

I want to be careful about what this work does and does not currently deliver, because overselling it would repeat exactly the error the piece is warning against. These are early results, demonstrated on small systems — a Mecanum bot, a pendulum — not on a bimanual manipulator doing contact-rich assembly. The guarantee in both cases runs from an approximate mathematical model to a high-fidelity simulator. That is a formalisation of one link in the chain. The remaining link, from the high-fidelity simulator to the physical machine on your floor, is the one the safety case ultimately needs and it is not yet closed by this line of work.

So this is not a technique to adopt next quarter for a production humanoid. The reason it matters now is the framing it establishes, and framings propagate faster than methods: a simulated result plus a bounded, stated gap is a fundamentally different artefact from a simulated result alone. The first is evidence. The second is a demo with excellent production values. Once a programme internalises that distinction, it starts producing better artefacts even with the tools it already has.

What the safety-side literature says is missing

The gap-quantification work approaches this from control theory. A parallel 2026 cross-layer analysis of safe embodied AI for long-horizon manipulation approaches it from the safety-engineering side, and lands somewhere uncomfortably compatible.

That work identifies four failure classes in long-horizon robotic manipulation — semantic misgrounding, subtask-level error propagation, execution drift, and contact-rich physical risk — and organises the available interventions by when they act: planning time, policy time, execution time. Its assessment of the state of the field is the part worth carrying into a design review. It reports limited evidence for policy-time safety, weak formal support for contact-rich manipulation, immature uncertainty-triggered intervention, and insufficient manipulation-specific safety benchmarks.

Read those two literatures together and the picture is coherent. The gap is largest where contact physics dominates. The formal tools are weakest in exactly the same place. And the benchmarks that would let you detect the problem empirically are, by the survey's own account, insufficient. Three independent lines of weakness converging on the same region of the operating envelope is not a coincidence; it is a description of where the risk actually lives.

The deliverable is a stated residual

All of which produces one practical consequence for anyone deploying under supervision, and it is deliberately narrow enough to act on.

The deliverable your safety case needs is not a higher-fidelity simulation. It is a stated residual.

If your validation says the policy achieved some result in simulation, the reviewer's next question — and it is the question that decides whether a later finding is contained or systemic — is: and what is your bound on how far the real system may deviate? A programme that can answer has an argument. A programme that cannot has a rendering.

What a validation artefact has to carry None of it is exotic. Almost none of it is currently produced. 1 Coverage stated What is modelled · what is approximated, and by what method · what is absent. 2 Residual quantified For the specific behaviours the safety case rests on — not an aggregate success rate. 3 Hardware testing aimed at the residual Designed to probe where the simulation is weakest, not to re-confirm the happy path. 4 Attestation split by source Which claims came from simulation, which from hardware, under which platform version. A solver change is a revalidation trigger. Treat it like a silent model update. 5 Gap growth monitored Hardware wears; environments drift. The residual is not a constant. Acceptable at commissioning is not automatically acceptable in year three. The reviewer's question is not what did the simulation show. It is: what is your bound on how far the real system may deviate? A programme that can answer has an argument. One that cannot has a rendering.

None of what follows is exotic, and none of it requires the certified-transfer research to mature first. State the simulation's coverage explicitly: what physics is modelled, what is approximated and by what method, what is absent. Quantify the residual for the specific behaviours the safety case rests on rather than in aggregate — an overall success rate hides precisely the concentrated, contact-adjacent failures that matter. Design hardware-in-the-loop validation to target the residual rather than to re-confirm the happy path, which is what most physical campaigns following a successful simulation actually do.

Then two things that are about time rather than about the initial validation. Attest what was established in simulation, what was established on hardware, and which claims rest on which — including the platform and version the simulation ran under, because a solver change is a revalidation trigger in exactly the way a silent model update is for a model-risk team. And monitor for gap growth: hardware wears, environments drift, and a residual that was acceptable at commissioning is not automatically acceptable in year three.

What domain randomisation actually buys, and what it quietly assumes

Domain randomisation deserves its own treatment, because it is the single most widely deployed answer to the reality gap and because the way it is usually described obscures the assumption it rests on.

The technique is elegant. Rather than trying to model the real world exactly — a losing battle — you randomise the simulator's parameters across a range: friction coefficients, masses, latencies, lighting, textures, sensor noise. A policy trained across that spread cannot overfit to any single parameterisation, so it learns behaviour that works across the whole range. If the real world falls inside the range, the real world is just one more sample the policy has effectively already seen. It works, repeatedly and across domains, and it is one of the genuine successes of the field.

Now notice what has happened to the assumption. Before randomisation, the programme assumed the simulator was accurate. After randomisation, it assumes the randomisation range contains reality. That is a better assumption — much better, because a range is more forgiving than a point — but it is still an unverified assumption about the world, and it has been moved somewhere less visible. Nobody writes “we assume friction on the deployed line falls within the sampled interval” on the validation summary. The interval is a training hyperparameter, buried in a config, chosen by an engineer using judgement.

Two consequences follow that matter for a safety case. The first is that the range is auditable in principle and almost never audited in practice, even though it is a load-bearing assumption of the entire validation. Asking to see the randomisation ranges, and asking what evidence supports them as covering the deployment environment, is one of the highest-yield questions available to a reviewer and costs nothing.

The second is that randomisation trades peak performance for robustness. A policy trained across a wide range is typically worse on any specific configuration than a policy tuned for it. That trade is usually worth making, but it means “widen the ranges” is not a free action when a residual concern is raised — widening degrades the policy, and past some point degrades it below usefulness. The technique has an internal limit, and programmes tend to discover it under deadline pressure rather than design for it.

Four failure classes, and what each does to a safety case

The cross-layer analysis of long-horizon manipulation names four failure classes. They are worth walking individually, because each degrades a different part of a safety argument and they are commonly discussed as if they were one problem called “the model made a mistake.”

Semantic misgrounding. The system's understanding of the task does not match what the task actually is — the referent is wrong, the object is misidentified, the instruction is interpreted against the wrong frame. For a safety case this is the most awkward class, because the system may execute flawlessly. Every metric a validation tracks can look excellent while the machine competently does the wrong thing. Performance monitoring will not catch it; only something that checks the grounding itself will.

Subtask-level error propagation. In a long-horizon task, a small error early becomes a large error late, because each subtask's starting state is the previous subtask's ending state. This breaks per-step validation as a methodology: a policy with excellent per-subtask success can have poor end-to-end success, and a safety case built on step-level metrics will overstate the system's reliability by a margin that grows with horizon length.

Execution drift. Behaviour deviates from intent during operation, without any discrete failure event to trigger an alarm. This is the class most directly coupled to the reality gap, because drift is what a residual looks like when it accumulates over time rather than appearing at a single decision point. It is also the class most likely to be invisible in simulation, where the drift sources — wear, thermal effects, cumulative calibration error — are frequently not modelled at all.

Contact-rich physical risk. The danger inherent to manipulation itself: forces applied, objects moved, contact made. This is where the consequences live, and by the survey's own assessment it is where formal support is weakest and manipulation-specific safety benchmarks are least adequate.

The compounding is the point. The gap is largest where contact dominates. The formal tools are weakest in the same region. The benchmarks that would surface the problem empirically are, by the field's own account, insufficient. Three independent weaknesses converging on one part of the operating envelope is not coincidence — it is a map of where the risk actually sits, and it argues for putting the physical testing budget there rather than spreading it evenly.

Who is going to ask, and on what authority

It is fair to ask whether any of this will be demanded, or whether it is a consultant's counsel of perfection. The honest answer is that the demand is not yet formalised anywhere, and the reason to act early is structural rather than regulatory.

Start with what is actually on the books. The machinery-safety and functional-safety frameworks — IEC 61508, ISO 13849, and for robots specifically ISO 10218 — govern the cell and are mature, but they were written for systems whose behaviour is specified deterministically, and they do not tell you what to do when the thing inside the envelope is a learned policy. They constrain the envelope well. They are largely silent on the occupant.

The model-risk world offers the closer analogy, and it comes with a warning attached. US supervisory guidance on model risk was substantially revised in April 2026 — OCC Bulletin 2026-13 and its parallel Federal Reserve and FDIC issuances — superseding the older framework that practitioners still reflexively cite. Two features of that revision matter here. It states that generative and agentic AI models are novel and rapidly evolving and are therefore not within its scope, with separate guidance promised. And it explicitly does not set out enforceable standards or prescriptive requirements.

It would be a serious misreading to treat that as an exemption, and the same misreading is available in robotics. Every obligation attached to the underlying activity — safety and soundness, worker safety, product liability, sectoral duties — is untouched. What was removed is the framework that would have specified the controls. Nobody is going to hand a physical-AI programme a control specification, and when the separate guidance is eventually written, it will be written against whatever the industry has by then built. That is an argument for building the evidence layer now and being one of the programmes the specification is written around, not an argument for waiting.

In the meantime the pressure will arrive from parties who do not need a regulation to ask a hard question. Insurers pricing a line. Acquirers running diligence on an automated facility. A customer's own risk function. And, most reliably, opposing counsel after an incident — who will ask what was validated, how, and on what basis anyone believed it transferred. None of those parties are bound by the pace of rulemaking.

What a stated residual looks like on paper

Abstract advice invites abstract compliance, so here is the shape of the artefact, for a single behaviour rather than a whole system — which is itself the first discipline.

Name the behaviour narrowly: not “the picking policy is safe” but something like “during the place phase, peak contact force on a compliant object stays under the specified limit.” Then state the coverage: which physics the simulator models directly, which it approximates and by what method, and what it does not represent at all — a list that usually surprises the people who commissioned the simulation.

State the randomisation ranges used for the parameters this behaviour depends on, and say what evidence supports those ranges as covering the deployment environment — measurements from the actual line, vendor specifications, or, honestly, engineering judgement, which is acceptable if it is labelled as such. Then give the residual: the measured or bounded discrepancy between simulated and physical outcomes for this behaviour, from whatever hardware trials exist, with the sample size and conditions attached.

Then the two temporal items. What is the trigger for revalidation — a platform or solver version change, a fixture change, a policy update, a wear threshold. And what is monitored in production to detect the residual growing, with the threshold at which the behaviour stops being trusted for autonomous execution.

A page like that, for the handful of behaviours a safety case actually rests on, is worth more than ten million simulated hours reported in aggregate. It is also, in my experience, the document that changes the engineering: teams that have to write down what is absent from the simulation start noticing things they had stopped seeing.

Three markets, three different starting positions

This is not a Silicon Valley story, and the right first move differs sharply by where the deployment sits.

In North America, the machinery-safety and workplace-safety framework already governs the cell and the institutional muscle for functional safety exists. The open question is narrower and harder: who accepts responsibility when the thing inside the cell learns, and how the existing performance-level apparatus is supposed to consume a statistically characterised policy. The work here is integration with a mature regime, not creation of one.

In India, an industrial build-out is deploying automation at speed alongside a modernising factory-safety regime. That timing is a genuine advantage if it is used: the assurance layer can go in with the automation rather than being retrofitted a decade later, which is the more expensive path every industrialised economy has already taken. The constraint is capacity — the number of people who can write a credible safety case for a learned system is small everywhere and smaller here, which makes the documentation discipline above more valuable, not less, because it is transferable.

In the Gulf, giga-projects, automated ports and new-build logistics are being commissioned from a standing start. That is the single best position anyone can occupy. Designing the evidence layer in costs a fraction of adding it later, and there is no legacy in the way. The risk specific to this position is different: new-build programmes buy integrated stacks, and an integrated stack is exactly where the platform-dependency question below becomes acute.

The dependency nobody has named

One further consequence deserves stating, because it follows directly from taking simulated evidence seriously.

If your safety evidence is produced inside a simulation environment, then that environment's physics assumptions, solver behaviour, asset fidelity and version cadence are all inputs to your safety case. This is an ordinary supply-chain-of-evidence question, of the kind industrial risk functions name routinely in every other domain, and it is not yet being asked crisply about simulation platforms.

The concrete version: what happens to your evidence when the platform ships a solver change? Do you revalidate? Would you know you needed to? A model-risk team handles the analogous problem with version pinning and change triggers, because they learned the hard way what a silent update does to a validated system. The physical-AI equivalent is to record the platform and version every validation ran under and to treat platform changes as revalidation triggers — which requires, first, that someone is tracking them.

Where this argument is weakest

Three places, stated plainly, because a piece arguing for stated residuals should carry one.

First, “quantify the residual” is much easier to write than to do. For a contact-rich manipulation policy there is currently no accepted method that produces a defensible number, which is exactly what the cross-layer analysis means by weak formal support. A programme that takes this advice seriously in 2026 will produce a bound that is partly qualitative, partly empirical, and honestly caveated — which is still a large improvement on no bound, but is not the clean artefact the argument gestures at.

Second, the certified-transfer work is at a scale far below industrial deployment and covers a different link in the chain than the one the safety case needs. I have cited it for its framing rather than its readiness, and that distinction should not blur.

Third, none of this has been tested against an actual regulatory finding. The argument that reviewers will ask the residual question is a prediction about how supervision will evolve, grounded in how it evolved for model risk in financial services, not an observation of enforcement that has already happened. It may arrive slower than I expect. What makes it worth acting on early is the asymmetry: a programme that can state its residual loses very little if the question never comes, and a programme that cannot loses a great deal if it does.

The sentence that carries across

The physical-AI stack is arriving faster than the assurance practice around it, exactly as the agentic software stack did over the past two years. The same sentence turns out to be true in both worlds. In software, an agent that acts inherits the obligations of the human it replaced. In the physical world a machine that moves inherits them too — and it inherits mass, momentum and proximity to people along with them.

Simulation is how you find out whether it works. It is not, by itself, how you prove it.