Start where the teardown ended, with a sentence that sounds like a technicality and is not.
The credential lifecycle discipline in this sector is real. It says who may hold access, on what basis, reviewed how often, and revoked when. Every clock inside it starts on the same class of event: a person is terminated, transferred, or has their access reviewed; an interactive session ends. Those events are observable, they are already wired into systems that exist, and the discipline built on them works about as well as such things work anywhere.
An agent's credential is owned by a field in a configuration management database. The field says a name. The field was correct on the day someone typed it. A field cannot be terminated, cannot transfer, and does not have an interactive session that ends — so none of the events fire, no clock starts, and the credential persists for as long as the system it authenticates to persists. Which in this sector is measured in decades.
The construction below does one thing and then builds three things on top of it. The one thing is manufacturing an event.
The gate: mint against a role, not a field
The design's first move is a refusal, and it happens before anything is issued.
A credential request arrives from a platform — per deployment, per agent, per connector. The gate asks two questions and refuses on either.
Who is the owner of record? A named natural person. Not a team, not a mailbox, not a vendor company, not a distribution list. Each of those four is a way of writing down an owner such that nobody is the owner, and each of the four is what estates actually contain. The gate refuses rather than warning, because a warning in an issuance path is a thing that gets acknowledged by whoever is trying to get their work done.
Do they currently hold an operational role? A control room position, an area authority, a maintenance planner, a named vendor sponsor — confirmed against the roster or the permit system rather than asserted in the request. The credential is then minted carrying the role identifier, not merely the person's name.
That last detail is the whole mechanism. A name is a string and decays silently. A role is a thing the organisation already tracks, already fills, already vacates, and already has a process around vacating. So the moment the role is vacated, an event exists — and the event is the thing the sector's existing lifecycle clocks are built to consume.
Vacancy opens a review window rather than triggering an immediate revocation, and the distinction matters operationally. A shift supervisor changing roles should not silently invalidate the credential a predictive-maintenance agent has been using for eight months; it should raise a question with a deadline attached. Non-renewal at the end of the window is what produces expiry. The design's claim is not that vacancy is dangerous. It is that vacancy is observable, and that the estate today observes nothing.
The right-hand column of Figure 1 is not a straw man. An owner field in a configuration database that is correct at write time and decays thereafter is the normal state of a well-run estate, not a badly run one. The failure is not carelessness; it is that nothing in the surrounding machinery is capable of noticing the decay, so noticing has to be somebody's discretionary effort — and discretionary effort is what gets deferred during a turnaround.
The calendar: rotation that can actually be executed
Give an estate the gate above and the next thing it needs is an expiry that means something. This is where most credential policy in this sector dies, and it dies for a reason that has nothing to do with security engineering.
Fixed-interval rotation — ninety days, pick your number — lands wherever the calendar puts it, which means it lands mid-run. Two things then happen, and both are bad. Either the rotation is skipped, because the change regime refuses a change to a running plant and refuses it correctly; or the rotation is executed and is hazardous, because changing a credential can break a data path that something downstream depends on, and the something downstream is frequently a historian feeding a model feeding a decision.
So the credential's clock becomes the plant's clock. Credential lifetimes are set so that each terminates inside a scheduled outage window, with a grace band before the window in which renewal opens. Rotation stops being an interruption and becomes a scheduled task with an owner, competing for the window alongside every other scheduled task — which is the form of work this sector is extremely good at planning and extremely bad at accommodating when it arrives unplanned.
The lower lane in that figure is the part I would have left out of an earlier version of this design, and leaving it out would have made the design useless. There is a population of credentials that cannot be re-credentialled within any window: hard-coded at commissioning into equipment whose vendor no longer exists, embedded in a historian's configuration that nobody has the source for, shared across a control system whose supplier's support agreement ended during a previous decade. A rotation design that has no lane for these is a design for a plant that does not exist.
The lane works like this. Each such credential carries a declared compensating control — narrowed reachability so the credential can only reach what it actually needs, session recording, an alerting rule on use outside an expected pattern, or a broker placed in front of it so the credential itself is never handed to the caller. And then the important part: the expiry is placed on the declaration rather than on the credential. The exception has to be re-argued on a clock, by a named person, rather than persisting quietly because the credential it protects cannot expire.
A rotation policy that cannot be executed is not a control. A declared exception with a clock on it is.
The choreography: revocation is a manoeuvre
Everything above is preparation. The part that distinguishes this sector from every information-technology analogue is what happens when a credential actually has to be withdrawn.
In an IT estate, revocation is an API call. You take the decision, you make the call, the credential stops working, and the blast radius of getting the timing wrong is a failed job and an annoyed engineer. In an operating plant the same call can land in the middle of a procedure, and the procedure may be one that cannot be stopped where it is.
The decision arrives for one of four reasons — compromise suspected, role vacated, contract ended, scheduled expiry reached — and the first question is not about the credential at all. It is about the world: is this credential in use inside a procedure right now?
- Idle. Revoke immediately. Then verify at the resource that the credential no longer authenticates, and record the verification. The verification step is not ceremony — a revocation that was issued and did not propagate is the failure mode that looks exactly like success in every dashboard.
- In an interruptible procedure. Hold at the next defined safe step, confirm the safe state with the control room, then revoke and verify. The hold carries a maximum duration, after which the third path is taken instead — because an unbounded hold is a revocation that quietly never happens.
- In a procedure that cannot be safely interrupted. Contain rather than cut: narrow the credential's reachability to the procedure in progress, raise monitoring, notify the control room and the shift supervisor by name, and revoke on completion.
The third branch is the one that makes this a sector-specific design rather than a generic one, and it is also the one that will draw the sharpest objection, so let me state the objection myself: it means that in the case where you most suspect compromise, the design's default is to let the credential keep working for a while.
That is correct, and it is why the branch carries a hazard override. If the assessed risk of continued use exceeds the risk of interruption, the decision escalates to the operational authority and the credential is cut. Containment is a decision the operator owns, not a default the design imposes on them. The design's contribution is to make the trade-off explicit and to put it in front of the person whose job it is to make it, rather than resolving it silently in either direction — which is what both alternatives do. A security tool that cuts unconditionally has decided that interruption is always safer. A process with no branch at all has decided that continued use is always safer. Neither of those people is in the control room.
The three branches, the hold times and the escalation path in this section are a constructed design, assembled from patterns in published operational practice and industrial security guidance. They are not a description of any real plant, operator, incident or deployment, and no part of this piece describes client work. The specific durations a plant should use are a function of its procedures and belong to the people who own them.
The plane: enforceable with the link down
The band across the foot of Figure 3 is a separate requirement and it is the one most likely to be skipped, because it costs infrastructure and its absence is invisible until precisely the moment it matters.
A revocation decision that can only be enforced by reaching a corporate identity service is a revocation decision that fails during exactly the events that motivate revocations. An incident that compromises a credential is frequently an incident that involves the network — and the standard, correct, well-drilled response to a suspected network compromise is to sever links, including the one the site would need to ask whether a credential is still valid.
So the revocation state has to be present locally, and verifiable locally. A signed revocation list cached at the site, checkable without reaching anything upstream, carrying a defined maximum staleness so that a site running on a cached list knows how old its knowledge is and can act accordingly. The staleness bound is the honest part: it says that in a disconnected condition this design gives you a known-imperfect answer, and tells you exactly how imperfect.
The alternative that estates actually run is worse in a way that is easy to miss: a system that fails open during disconnection, so that a severed link silently converts every revoked credential back into a working one. That is not a hypothetical failure mode. It is the default behaviour of a great deal of infrastructure that treats an unreachable authority as an unknown and an unknown as a pass.
Paying the bill
Four costs, and one of them is a genuine new attack surface that this design creates and did not exist before it.
The roster becomes a security dependency. The gate confirms operational roles against the roster or permit system, which means those systems are now in the credential-issuance path. They were previously HR and operations systems with HR and operations availability expectations. If the roster is wrong, credentials are wrongly refused or wrongly granted; if the roster is down, issuance stops. That is a real coupling and it should be planned for rather than discovered.
Turnarounds slip, and now credentials slip with them. Snapping expiry to the maintenance calendar means a deferred outage defers a population of expiries in one movement. The mitigation is that the grace band before each window has to be wide enough to absorb ordinary slippage, and that a deferral past the band is itself an event requiring a declared exception — which puts it in the lane that has a clock on it. It is a cost, not a flaw, but a design that snapped expiry to the calendar and said nothing about slippage would be a design that had not been near a plant.
The compensating-control lane can become a permanent parking space. This is the failure I would bet on. The clock on the declaration is the countermeasure, but a clock only works if the re-argument is genuinely capable of failing — and in an organisation where the answer has been yes for three years running, the fourth review is a formality with a signature on it. I do not have a mechanism that fixes this. The best I have is that the declaration names an individual rather than a function, and that the population under declared exception is reported as a number that is supposed to go down.
The containment branch is a target. An adversary who understands this design has an incentive to ensure that any credential they compromise is always inside an uninterruptible procedure, because that is the branch that buys them time. The design does not solve this. What it does is make the branch a decision with a named owner and an active escalation rather than an automatic behaviour, so that repeatedly landing in it is visible to a person rather than absorbed by a rule. Treating the frequency of third-branch outcomes as a monitored signal, rather than as an operational statistic, is the countermeasure I would add — and I am not confident it is sufficient.
What would falsify this
Three things, stated so the design is answerable rather than merely arguable.
First: if an estate can demonstrate that its machine credentials already have a starting event — that something observable fires when the reason for a credential ends — then the gate in Figure 1 is solving a problem that estate does not have, and the rest of the construction can be built on whatever that event is instead. I have not seen it, but I have not seen every estate, and the teardown's finding was about the instruments rather than about every implementation of them.
Second: if credential rotation in a running plant turns out to be routinely safe — if the data-path breakage I have treated as the reason rotation gets skipped is rarer than I think — then the calendar snapping in Figure 2 buys much less than it costs, and fixed-interval rotation with good testing is the simpler answer. This is an empirical question and I would want it answered with incident data rather than with argument.
Third, and most seriously: the whole construction assumes the credential is the thing worth controlling. If an agent in this environment can reach an effect without presenting a credential at all — through a trust relationship established at commissioning, through a network position, through a protocol with no authentication in it, which describes a great deal of what is actually deployed — then a perfect credential lifecycle governs a door that has a wall missing next to it. The teardown named that estate honestly and this piece does not fix it. Nothing here should be read as a claim that credential discipline is sufficient. It is the part that is buildable now.
That is the honest boundary of a construction rather than a product. It closes one gap, it names what it does not close, and the reason to build it during this interval is the same as it was in the teardown: the instrument that will eventually specify these controls will be written by people reading what the sector built while nothing was specified.