On 29 July, NC AI and CMES Robotics signed a memorandum of understanding for joint physical-AI development. The arrangement is described in the announcement as a “virtuous-cycle development system”: CMES provides data accumulated in manufacturing and logistics sites, NC AI uses it to train, and trained models are applied back to robot systems.
Read that as an assurance problem rather than an engineering one and something uncomfortable appears.
CMES Robotics is a member of the K-Physical AI Alliance — the 53-organisation consortium NC AI launched in February to compete for a government-backed initiative on world and robotics foundation models. Within that alliance, CMES is one of the parties contributing to the robotics foundation model itself. It is also, per the MOU, the party supplying the live industrial data that trains it, and the party whose robot systems the trained models are applied back to.
The operational environments where alliance models are tested — manufacturing facilities, logistics centres, hotels, airports — are provided by Samsung SDS. Also a member.
Three roles that independent assurance would deliberately keep apart — building the model, supplying its training data, and providing the environment it is validated in — sit inside one consortium, whose members are competing together for the same mandate.
This is a structural description, not an accusation
I want to be precise, because the easy version of this argument is cheap and wrong.
Nothing here suggests anyone is acting improperly. Consortium structures are how capital-intensive industrial R&D actually gets funded, in every country that does it well. The 53 members include 15 research institutions and 38 demand-side partners — a deliberate attempt to keep the work anchored to real operational needs rather than to benchmarks, which is a good instinct and rarer than it should be.
The contribution split is specific and sensible: two members on the robotics foundation model, one on physics-based simulation, one supplying dual-arm platforms. Nobody is pretending to do everything. That is a healthier shape than most alliances manage.
The public description of this programme names no independent evaluator.
Not a hostile one. Not a regulator. Simply a party whose interest in the result is different from the interest of the parties producing it. That absence is not unusual — it is close to universal in this field right now, which is exactly why it is worth naming before the pattern sets.
What makes physical different
The evaluation field has spent years learning that a benchmark score obtained the wrong way sits in the results table looking exactly like a legitimate one. The number is not false — the model really did produce those outputs. It is unmoored, because the process that produced it was not the process the number is supposed to summarise.
The alliance has already named the physical version of this, and the term is better than anything I would have coined: physical hallucination — generative models failing to adhere to real-world physical laws. That naming is a real contribution. It is precise, and it puts the problem in the gap between what a model asserts and what the world will honour.
Here is what the naming does not resolve. A hallucinating language model produces a sentence that is wrong. A hallucinating world model produces an action that is wrong, executed by a machine with mass. The intended applications named for this class of system are semiconductor cleanroom operations, steel production processes and shipyard block assembly — domains where a wrong output is not a bad row in a spreadsheet.
A benchmark that is wrong costs you a claim. A world model that is wrong costs you whatever it was holding.
Stated as a property of the system class — what an actuator can do — not as an allegation about any operator's plant. Nothing in the public record describes harm at any facility named here, and nothing in this piece should be read as claiming otherwise.
The procurement shape, one industry over
Anyone who has bought containment rather than built it will recognise this. When your isolation boundary is procured, the escape suite behind it — where one exists at all — was almost certainly commissioned and run by the party with the strongest possible interest in a clean result. That is not an allegation about any vendor's integrity. It is a description of an incentive structure, and it is the entire reason independent adversarial testing exists as a category in every other safety-critical field.
Aviation did not arrive at independent certification because manufacturers were dishonest. It arrived there because “we tested it ourselves and it passed” carries less information than it appears to, no matter who says it, and the consequences of being wrong were measured in people.
Physical AI is heading into the same territory with the same structure and, so far, without the same correction.
What the claim would need to mean something
Four declarations turn a fidelity claim from an assertion into a result. They transpose directly from evaluation to physical systems, and they cost nothing but the discipline of writing down what you actually did.
Who tried to make it fail, and from where. A validation run by the team that built the model, in an environment supplied by a partner, with a mandate not to damage anything, is a fundamentally different artifact from one run by an external party with permission to be destructive. Both are useful. They are not the same claim and should not produce the same sentence.
The denominator. How many trials, under what distribution of conditions? “No failures observed” without a count is a statement about the observer. Zero from three runs is not evidence. Zero from three thousand, with the conditions declared, is a result.
The conditions not tried. The most valuable line in any validation report and the one most often absent. Which lighting, which surface friction, which sensor degradation, which human-proximity cases were explicitly out of scope? A report that says nothing implies coverage of everything, which no physical test achieves.
An expiry. A fidelity result describes a plant on a date. Plants drift — a fixture is replaced, a camera is remounted two centimetres left, a line speed is raised for a quarter. A twin whose claim has no expiry is asserting a property of a factory that no longer exists.
The sim-to-real gap is usually discussed as a modelling problem — the model trained at one resolution meets a camera at another, the simulated friction was not the real friction. That framing is correct and incomplete. The gap is also an evidence problem: it persists partly because the conditions under which a policy was validated are so rarely declared that nobody can tell which gap they are looking at.
Three objections that deserve answers
“Independent evaluation for physical AI does not exist yet, so this demands something unavailable.” Largely true, and the strongest objection. There is no mature third-party escape-suite equivalent for world models acting on industrial hardware. But the absence of a mature market is an argument for building the capability, not for treating self-validation as equivalent. The four declarations cost nothing and can be adopted unilaterally this week by any party willing to state its method.
“Consortium structure is how this gets funded at all — you are asking for a purity that would stop the work.” Also fair. Government-backed consortia exist because no single firm can carry foundation-model development against industrial hardware. The argument is not against the structure. It is that the structure creates a specific epistemic gap, and the gap should be named in the programme's own documentation rather than discovered by whoever is standing near the equipment.
“The model already reports strong efficiency at high task success.” That is the builder's own reported figure from its own demonstration, and it is a claim about efficiency — exactly the class a builder can credibly self-report, because resource consumption is measurable on your own hardware. Physical safety under adversarial conditions is not the same kind of claim and does not inherit the credibility.
What to ask, if you are buying any of this
- Who validated this, and what is their relationship to whoever built it? If the answer is “a partner,” you have a design statement, not a test result. Record it as one.
- What is the denominator, and under what conditions? Trials, distribution, duration.
- Which failure conditions were explicitly out of scope? If nobody can answer, the scope was never bounded, which means it was never defined.
- When does the fidelity claim expire, and what invalidates it? Name the plant changes that would void it.
- What happens to the evidence if the relationship ends? Who holds the validation record, and can you read it?
A supplier who answers all five well is doing something genuinely rare and should be told so. One who cannot answer the first is not selling validated autonomy — they are selling capable software, which may still be exactly what you want, provided it goes in the risk register under its real name.
What this cannot tell you
- Whether these particular models are safe. I have no access to their validation record and make no claim about it. The argument is about what the public description does and does not contain.
- Whether the loop is closed at facility level. The announcement says trained models are applied back to robot systems. It does not say they return to the same sites that produced the data, and I checked specifically because the sharper version of this story would have depended on it.
- Whether this approach is worse than anyone else's. It is more legible, because more of it has been published. Most programmes elsewhere have not described their structure this openly, and legibility should not be punished with harsher scrutiny than opacity receives.
- Whether any of this slows deployment. The declarations proposed here are documentation of work already being done.
The uncomfortable summary is that the hardest problem in physical AI right now is not the sim-to-real gap. It is that we have not built the institution that would tell us how wide the gap is — and the parties currently best placed to measure it are the ones with the least interest in a large number.
They named the failure mode. Naming the referee is the harder half, and it is still open.