Ask an engineering team what their stop control does and you will usually get an answer about mechanism: it revokes the credential, it halts the workflow, it trips the interlock. Ask instead what it establishes — what you know to be true once you have pulled it — and the conversation gets slower.
That second question is the one that matters, and the answer depends entirely on which side of a line the effect sits.
The word stop is doing two jobs
On the information side, stop means: no further effects, and the effects that already happened can be brought back to where they were. A record restores. A message is retracted or corrected. A transaction is reversed by a compensating entry. The cost of stopping late is real — it is downtime, it is effort, sometimes it is a great deal of both — but the state is recoverable, and every control designed for that world has that assumption underneath it.
On the physical side there is no such assumption available. A hole has been drilled. A weld has been laid. A batch has been released to the next stage. A valve has been opened, a dose delivered, a breaker operated. None of these has an inverse operation. You can do further things afterwards — plug the hole, cut out the weld, quarantine the batch, close the valve, reclose the breaker — and none of them is an undo. They are new actions with their own consequences, performed on a world that now contains the first action permanently.
So on the physical side the question is never whether the effect can be reversed. It is whether the system can be brought to a state that somebody has previously characterised as safe — and whether anyone can show that it got there.
The path, link by link
Five links. Each has a failure mode that is specific, common, and mostly invisible until the day it matters.
One — the decision is taken. By a named person, or automatically by an operating condition leaving its envelope. It fails when nobody present is authorised to take it. A stop control whose authority sits with a role that is not staffed at three in the morning is a stop control with a callback in front of it, and the callback is measured in the units that matter here: minutes during which the thing continues.
Two — the decision reaches an enforcement point. Over a path that has to survive the incident that motivated it. This is where a great deal of otherwise careful design fails, and it fails for a reason that reads as almost unfair: the standard, correct, well-drilled response to a suspected compromise is to sever network links, and a stop decision that travels over a corporate link is a decision that arrives after somebody has correctly done their job.
Three — the enforcement point acts. Withdrawing the capability, or refusing the operation at the point of work. It fails when the agent already holds what it needs and never has to ask again. This is the argument for issuing capabilities per action rather than provisioning them once: a credential obtained an hour ago is a decision that was made an hour ago and cannot be unmade by anything happening now. Long-lived credentials in this environment are not a hygiene problem, they are a stopping problem.
Four — the system moves to a defined safe state. Defined per equipment and per operating condition, never globally. It fails when no safe state was defined for the condition the system is actually in, so it halts wherever the instruction found it — and in several equipment classes, stopping in an arbitrary position is more dangerous than continuing to a defined stopping point. "Safe state" is not a synonym for "off."
Five — the safe state is confirmed by measurement. From a source the agent did not produce. This is the link that fails silently, and it is the subject of the rest of this piece.
The fifth link is where the belief gets formed
Links one through four can all succeed while the system is still unsafe. The decision was taken, it arrived, the enforcement point acted, and the equipment did something — and none of that is a statement about the state of the world. It is a statement about the state of the control system's intentions.
What operators actually see, in almost every system I have looked at, is a confirmation that the command was sent. That confirmation is then read as a confirmation that the system is safe, because it appears at the moment the operator is asking whether the system is safe, and because it is green.
"Command sent" and "system safe" are separated by every physical thing that could have gone wrong. A valve commanded shut and a valve that is shut are different facts, and the gap between them is filled by actuator faults, mechanical binding, a positioner that reports its own command back to you, and a control loop that is doing exactly what it was told by something that was wrong. None of these is exotic. All of them are the ordinary reasons process industries built independent measurement in the first place.
So the requirement is specific: the state that confirms the safe state must come from a source that is not the thing being stopped, and not the agent that was acting. If the agent's own reporting is the evidence that the agent has been stopped, the evidence is a self-report from the subject of the investigation.
That principle is not new and this sector did not need me to discover it — independent measurement, diverse instrumentation and separated protection layers are old ideas here, learned expensively. What is new is that a machine principal is now in the loop at machine rate, and the confirmation link is the one that nobody has re-derived for it. The protection layers were designed on the assumption that the thing being confirmed was a human action or a control loop, not an autonomous process with its own view of what it did.
A stop nobody will pull is not a control
There is a failure mode that sits above all five links, and it is the one that quietly disables more stop controls than any technical defect.
If the only available stop takes down more than the thing that needs stopping, then pulling it has a cost that grows with the size of the estate — and at some size, the cost of pulling it exceeds the cost of the thing you are worried about, at which point the control still exists, is still tested, and will never be used. Everyone involved will make locally reasonable decisions and the outcome is a control that is decorative.
So scope is a safety property, not an ergonomics one. The ability to halt one workflow, on one piece of equipment, without halting the estate is what determines whether the stop gets pulled at the moment it should be. This is also, and not coincidentally, the thing two supervisors in different jurisdictions have independently written down as something a board should be able to evidence — proposed in India's draft model-risk guidance and enumerated in the Central Bank of the UAE's guidance note, both of which are draft or guidance-note status and neither of which binds anyone discussed here.
The convergence is worth noticing precisely because those instruments were not written for plants. They were written about AI systems generally, by people reasoning about accountability, and they arrived at the same requirement that process safety arrives at from the other direction. When two traditions with different vocabularies converge on the same control, the control is telling you something about the problem.
Everything in this piece that describes an operating scene, a failure sequence or a stopping decision is a constructed illustration assembled from published operational practice and general engineering literature. None of it reports a real plant, operator, incident or deployment, and no part of it describes client work. The sector's functional-safety and industrial-security series are referenced at series level only; no individual standard is cited, and nothing here claims that any instrument requires this construction.
Paying the bill
A kill path is not free and it is not purely protective. Three costs, and the second is a genuine new attack surface.
It is an availability risk. Every mechanism capable of stopping production is a mechanism capable of stopping production by mistake. A condition-triggered stop with a badly chosen envelope will fire during ordinary operation, and it will do so at three in the morning, and the organisational response to that is to widen the envelope until it stops firing — which is how a safety control gets tuned into inactivity by people acting reasonably.
It is an attack surface. An adversary who cannot reach the process may be perfectly satisfied to reach the stop. A kill path with weak authentication is a denial-of-service capability with an operator-friendly interface, and the more decentralised you make it for availability reasons, the more places it can be reached from. I do not think this argues against building one. I think it argues that the kill path deserves the same authority discipline as the thing it stops, which is a sentence that costs a great deal more than it reads.
Defining safe states per condition is expensive and nobody wants to fund it. The engineering work is enumerating operating conditions and characterising a safe state for each, which is unglamorous, requires the people who are busiest, and produces an artefact whose value is invisible until an incident. Every version of this design I can construct requires that work, and I have no way to make it cheap. What I can say is that a design which skips it has not eliminated the cost — it has deferred it to the moment when someone is deciding under time pressure what safe means for this equipment in this condition.
What would falsify this
Three things, and I would want to hear any of them.
First: if independent confirmation of the safe state turns out to be routinely available already — if the fifth link is in fact the well-built one in most estates and my sense that it is the weak one is an artefact of the systems I have looked at — then the emphasis in this piece is misplaced, and the real work is at links two and three. This is an empirical question about deployed practice and I have a small sample.
Second, and more fundamentally: this whole construction assumes the capability is the thing that reaches the effect. If an agent in a given environment can cause a physical action without presenting anything a control can withdraw — through a trust relationship established at commissioning, through a network position, through a protocol with no authentication in it — then a perfect kill path governs a door standing next to a gap in the wall. That describes a meaningful fraction of what is actually deployed, and nothing in this piece fixes it.
Third: the argument treats reversibility as a property of the effect, and there are edge cases where that is too crude. Some physical effects are partially reversible at meaningful cost, and a design that classes them with the irreversible ones will over-engineer them. I have drawn the line where I have because the failures I care about are unambiguous, but the line is coarser than reality and someone working in a specific equipment class will know better than I do where it actually sits.
The reason to publish the construction rather than the conclusion is that each link's failure mode is checkable against a real system by the people who own it. If your kill path has all five links, and the fifth one measures from somewhere the agent cannot reach, then this piece is describing something you have already built — and I would rather learn that than be right about the gap.