Take the constructed illustration the teardown ended on and change one thing.

A transmission control centre on a July afternoon. The state estimator solves, a contingency screen flags a post-contingency thermal overload on a corridor, and an advisory agent drafts a switching sequence from the pre-studied case library. Eight months earlier an operator would have read it and clicked. Now an auto-execute rule handles the routine class, and the command goes to the gateway and out to the substation without anyone touching anything. The change that created the rule was properly approved. The access was properly provisioned. Nothing was hidden.

Change this: before the command reaches the gateway, the gateway asks a question. Which grant is this being executed under. It resolves an identifier, finds a grant issued at 06:00 by a named certified operator at the start of their shift, covering this study class and this named apparatus, valid while the area is N-1 secure and no major incident is declared. It then takes its own reading of the system — from the state source it holds, not from the agent that produced the proposal — and checks the predicate. The predicate holds. The command goes through, and the record that lands in the historian names the grant, the operator, their certification, the conditions asserted at 06:00, and the conditions measured at 14:31:07.

Eight months later, when the post-event review asks which certified operator authorised the switching action that opened that breaker, the question has an answer. Not a better-written ticket. An answer, of the type the question asked for, produced by the same object that was constraining the machine at the time.

The control room, the shift, the grant, the timestamps and every operational detail above and below are a constructed illustration, assembled from patterns in public specifications, published reliability guidance and vendor documentation for energy management systems. It is not a report of any real event, utility, operator or deployment, and no part of this piece describes client work.

That is the whole construction, and the rest of this piece is why each part of it is shaped the way it is — starting with the two objections that have to be conceded properly before any of it earns a hearing, because both come from people who know this sector better than any generalist writing about AI.

Conceding the change regime, properly

The first objection is that this sector already controls who may change what, more rigorously than almost any other, and that describing an agent as acting unauthorised inside a utility is describing a utility that does not exist.

It deserves the full version. Switching on a transmission or distribution system runs on switching orders: written, reviewed before execution, executed step by step with verification at each step, with the operator repeating back the instruction and confirming the operation. Material changes to the energy management system travel as change requests with test evidence, back-out plans and named approvers, through a change board, into a scheduled window. The EMS itself is a controlled system: the vendor's releases are regression-tested, the network model is under configuration management with its own change discipline, and the state estimator's model updates are treated as changes because everybody has seen what a bad model does to a solution. Factory and site acceptance testing precede anything reaching production. The shift log is a legal-quality record kept by people who know it will be read after an event. Two-person verification exists for high-consequence operations, not as a policy aspiration but as the thing that actually happens on the bridge. None of this is decorative, and all of it was learned expensively.

The second objection is the CIP family, and it deserves the same treatment. Utilities operating bulk electric system assets run their access and change controls under mandatory, enforceable reliability standards with per-violation penalties — a compliance posture most sectors' AI conversations would envy. Electronic access to the critical systems is controlled at a defined perimeter, monitored, logged and revoked on defined clocks. Personnel with access are subject to risk assessment and to training obligations. Configuration change management on those systems is a compliance obligation with evidence requirements attached, and the whole apparatus is audited by regional entities that have no interest in being generous. I am naming the family and not a standard number deliberately: this piece does not need one, and a reliability reader will check any number that appears.

Both of those are true, and neither of them is the question. Here is where each one lands.

The change regime governs a change. The autonomy is not a change. A change request describes an intended action: this setting, this apparatus, this window, this back-out. The auto-execute rule was a change, and it was properly approved — and what it approved into existence was a standing capability to take a class of actions indefinitely, which the change discipline then never looks at again. Each subsequent action is not itself a change request; it is an instance of the approved automation. So the regime's controls attach at the moment the capability was created and at no moment thereafter, which is precisely inverted from what the sector does for humans, where the switching order is issued per operation and the authority is re-established every time.

The CIP family governs the perimeter and the access regime. It answers who may reach the system. That is a different question from which certified person authorised this action under these conditions, and the difference is not a gap in the standards — it is their scope. The family was written to keep unauthorised parties out of the systems and to evidence that access is controlled, monitored and revoked. It does that. Full compliance with every access and change requirement in the family is entirely consistent with the authority question having no answer at all, because the family's subject is reach and the question's subject is decision. Figure 3 later in this piece makes the point in its own terms: both the record that exists today and the record the construction produces comply with the family equally.

The switching-order discipline is not an objection. It is the design. The strongest thing about this sector is that it already solved this problem, in hardware, for people. A switching order is a written, bounded, condition-laden authorisation from a qualified person, issued before the action and verified at the point of work. Every property the construction below needs is in that sentence. The industry did not arrive at it theoretically; it arrived at it because the alternative killed people. What has not happened is its transposition to the layer where the machine now acts — and the reason is not reluctance, it is that nobody has offered a shape for it that runs at machine rate.

The boundary is a line between credential families

Start with the part that does the most work for the least cleverness.

In the shipping pattern, recommend and execute are two stages of one workflow inside one system running as one principal, and the boundary between them is a step: the operator's acceptance. Steps can be removed. The teardown's whole opening turns on a step being removed by a properly approved efficiency change, without anything about the principal's capability changing at all — because the capability never encoded the step in the first place.

So do not build the boundary as a step. Build it as a line between two families of credential, and put no principal on both sides of it.

  1. The advisory family. Read the state estimate and telemetry. Run contingency and power-flow studies. Read the outage schedule and the pre-studied case library. Write a proposal into the recommendation queue. That is the complete set, and every one of them is a read or a write into a queue that reaches no apparatus.
  2. The execution family. Issue supervisory control to the gateway. Operate apparatus through the RTU. Change protection and automation settings. Every credential in this family, without exception, reaches physics.
  3. The rule. No principal holds credentials from both families. The advisory agent holds the first set and nothing from the second. It cannot issue a command, not because a policy forbids it, but because it possesses nothing that a gateway would accept.
FIGURE 1 · THE CAPABILITY LINE A review step can be skipped. A credential the principal does not hold cannot be used. THE CAPABILITY LINE THE ADVISORY FAMILY · read the state estimate and telemetry · run contingency and power-flow studies · read the outage schedule and case library · write a proposal into the recommendation queue WHO MAY HOLD THEM the advisory agent · study and operations engineering · any analytics service no principal here holds supervisory control THE EXECUTION FAMILY · issue supervisory control to the gateway · open and close apparatus through the RTU · change protection and automation settings WHO MAY HOLD THEM certified system operators, on shift — and the gateway itself, acting as verifier no service principal appears in this column THE EXECUTION GRANT issued by a named certified operator scoped to a study class and named apparatus carrying the condition predicate it was issued against expiring on clock or on condition change · revocable proposal verified at the gateway The agent never holds execution authority. It holds a proposal. A person's grant is the only thing that turns a proposal into a command.

The difference between a step and a line is what happens under pressure. A step is removed by a reasonable person making a reasonable efficiency argument, and its removal is a configuration change that leaves the principal's capability untouched. A line is crossed only by provisioning an execution credential to an advisory principal, which is an identity change, in the identity system, with the identity system's own evidence attached — and it is the kind of change that a utility's access review would catch precisely because access review is the control this sector already does extremely well. The construction is not asking the operator to build a new discipline. It is arranging the problem so that the discipline they already have points at it.

What crosses the line, then, is not a principal and not a permission. It is a single object: a grant, issued by someone in the execution family, which authorises a specific class of proposals from a specific advisory principal to be executed under stated conditions. That object is the entire construction, and everything that follows is its anatomy.

Mapping the line to control-room roles

A capability line drawn in the abstract is an architecture diagram. It becomes operable when each side maps to a role that exists on a real shift roster, and the mapping has to be by what the role may grant rather than by what the role may do — because that is the only formulation that lets a certified operator authorise machine-rate action without being asked to occupy a path that moves faster than a person.

  1. The certified system operator on shift. Holds execution credentials directly, as today. Additionally, and this is the new authority, may issue autonomy grants within their certification class, bounded by their own shift, over apparatus within their area of responsibility. They cannot grant what they could not do.
  2. The shift supervisor. May issue grants that span a shift boundary, and may revoke any grant in the area without a handover conversation. Revocation is deliberately the easier privilege of the two: it should always be simpler to withdraw autonomy than to extend it.
  3. Operations engineering and the study function. Owns study classes — the families of pre-studied cases a grant can refer to. They do not issue grants and hold no execution credential; what they own is the definition of what a class means, which is where the engineering judgement actually lives.
  4. The outage coordinator. Owns the planned-works picture that feeds the condition predicate. Not a grant issuer either, but the source of one of the conditions that closes windows, and therefore a party the design has to name so that the condition has an owner.
  5. The advisory agent, and every analytics service beside it. Holds no grant-issuing capability and no execution credential. Its ceiling is a proposal.

The property that makes this work is that grant-issuing authority is bounded by the issuer's own authority, in both class and scope. An operator certified for one class of decision cannot issue a grant covering another. This is the ordinary principle that a delegation must be no wider than what the delegator holds, and it is worth stating explicitly here because the sector's terminal object is not a person but a credentialed person — so the chain has to check a certification, live, at issuance, and refuse to issue against a lapsed one.

This is the property the financial-sector version of the argument never needed. In banking, the requirement is that the chain terminates in a named natural person; any competent person will do, and the accountability regime supplies the rest. Here the terminal link has a type. A grant rooted in a named person who does not hold a live certification for that class of decision is not a weak grant — it is the wrong type of object, and a verifier should treat it the way a compiler treats a type error rather than the way a risk framework treats a finding. Building the certification check into issuance, rather than into a quarterly attestation, is the difference between the design respecting the sector's accountability model and merely gesturing at it.

The grant: window-scoped, condition-bound, twice-exiting

Now the object itself. Seven fields, two exits, and one property that decides whether the whole thing is real.

FIGURE 2 · ANATOMY OF A WINDOW-SCOPED GRANT The decision is pre-positioned. The record is not an afterthought — it is the same object. 1 · ISSUANCE a certified operator, ahead of need at shift start, or at study approval, or at window open — never in the action path, which has no room WHAT THE OPERATOR BINDS their identity and live certification the shift the grant belongs to the study class it covers the named apparatus in scope the conditions asserted at issuance this is a switching order, at machine rate THE GRANT issuer operator, by directory id certification class and expiry, checked live shift the seat, and its hours study class the pre-studied case family apparatus named, not domain-wide conditions the predicate, carried in full expiry clock — or condition change revocable by handle, without severing the account 2 · VERIFICATION at the gateway, before the RTU not in the agent's own stack, and not only at issuance WHAT THE GATEWAY CHECKS the grant resolves and is unexpired the class matches the proposal the apparatus is in scope the predicate still holds — measured from a source the agent did not produce a refusal is recorded as carefully as an execution expiry on clock end of shift, end of the issued window expiry on condition fires the moment the state leaves the predicate Both exits matter. A grant that expires only on the clock is an entitlement with a timer; the sector's own risk calculus was never purely temporal — the same action is routine under one system state and reckless under another. The decision lives ahead of the action because the action path has no room for it. Issued as an object, the evidence is a by-product.
  1. The issuer, by directory identity. A person, resolvable, not a role mailbox and not a team.
  2. The certification, checked live at issuance. Class and expiry. A grant cannot be issued against a certification that has lapsed, and it should not be issuable against one that expires inside the window it is granting.
  3. The shift. The seat and its hours. This is what ties the grant to a period during which someone is actually available to be asked, and it is the reason a grant should not routinely outlive a shift.
  4. The study class. The family of pre-studied cases the grant covers, owned by the study function. This is the field that carries the engineering, and it is the one that makes revocation surgical: withdraw a class and every grant referring to it stops, everywhere, without touching anything else.
  5. The apparatus scope. Named. Not the domain the credential could reach. The gap between those two is the same gap the telecom manifest surfaces in its own units, and it exists here for the same reason: supervisory-control credentials come scoped to a gateway, not to a breaker.
  6. The condition predicate, carried in full. Not a reference to a configuration item that can be edited afterwards. The conditions travel inside the grant, so that the grant is self-describing and a reviewer reading it eight months later does not have to reconstruct what the predicate said in July.
  7. The revocation handle. One call withdraws this grant, or every grant of this class, or every grant issued by this operator — without severing the service account and taking down every automated pathway including the ones behaving correctly.

Then the two exits, and the second one is the whole safety property.

Expiry on the clock is the exit everyone builds. It is not sufficient and it is not the interesting one. A grant that ends at the end of a shift is a good hygiene property and a poor safety property, because the sector's risk calculus was never purely temporal. The same action is routine under one system state and reckless under another, and a window that runs from 06:00 to 18:00 says nothing about whether the system entered a condition at 14:00 that the study never covered. A clock-only window is an entitlement with a timer on it, which is the thing the teardown identified and named.

Expiry on condition change is the exit that makes it a grant. The moment the measured system state leaves the predicate the grant is closed, whatever the clock says, and reopening it requires a person. This is what converts the operator's morning decision from an assertion about the afternoon into a constraint on the afternoon. It is also the property that makes the evidence honest: a record showing that a grant closed at 15:42 because the area left N-1 security, and that four subsequent proposals were refused, tells a reviewer something no telemetry archive can reconstruct — that the envelope was not merely described but enforced.

The property that decides whether any of this is real is where the predicate is evaluated, and by whom. It must be evaluated at the dispatch point — the gateway that hands commands to the field — using a state reading the agent did not produce. An implementation in which the agent evaluates its own predicate against its own view has built a system that decides for itself whether it is permitted to act, which is exactly the standing capability the construction was supposed to replace, wearing a new noun. This is the property most likely to be traded away during implementation, because preserving it requires a change in a team that does not own the agent, and it is the one I would refuse to trade.

Two consequences worth stating. A refusal is recorded as carefully as an execution, because a stream of refusals is the single best signal that the study class or the predicate is mis-specified — the control tells you where it is wrong, if you keep the evidence of it saying no. And the gateway needs a defined behaviour when its own state source is stale: refuse, and alarm. Fail-closed is the correct default at a boundary that touches apparatus, and an implementation that fails open when the state reading is late has inverted the entire design at the one moment it matters.

The gate as an evidenced grant rather than a console prompt

The teardown's sharpest observation about this sector is that a gate which has said yes four thousand consecutive times is a latency stage, and that the sector's own reliability body named the failure mode first: NERC's November 2024 white paper on AI and machine learning in real-time system operations is explicit that human-in-the-loop operation requires deliberate design if it is not to be human-in-the-loop in name only. The construction takes that seriously rather than arguing with it.

The response is not a better prompt. It is to stop treating the click as the control.

The click produces no object, which is why removing it changed nothing observable. In the shipping pattern, an operator's acceptance sets a flag and the plan proceeds. Nothing is created; nothing outlives the moment; nothing can be revoked, because there is nothing to revoke. That is why the efficiency change that removed the click was invisible in every artefact the utility keeps — not because anyone concealed it, but because the click's only trace was the latency it added. A control whose entire footprint is a delay is a control that can be deleted for performance reasons by someone acting in good faith.

A grant produces an object, and the object is what the gate was supposed to be for. When the operator's decision issues a grant, the decision has a lifetime, a scope, an owner, a set of conditions, and a handle. It can be audited without asking anyone what they remember. It can be revoked in one place. It can be counted — how many grants of this class were open last month, issued by whom, closed by clock or by condition — which turns the gate from an unobservable behaviour into a measured one. And crucially, it is the same object that enforces the envelope, so it cannot be quietly removed for latency, because removing it stops the automation rather than speeding it up.

There is an honest limit here, and it is the same one in a different place. The grant does not prevent decision decay; it relocates it. An operator who issues the same broad grant at 06:00 every morning without thinking about it has performed the same rubber stamp, one level up. What the construction buys is that the rubber stamp is now visible: identical grants issued daily with no variation in scope or conditions is a pattern anybody can query for, whereas identical clicks are invisible by construction. Relocating a failure from an unobservable place to an observable one is a real improvement and it is not a solution, and I would rather say that than claim the design closes a human-factors problem that no design closes.

What the reliability auditor sees

The last part of the construction is not architecture. It is the artefact a post-event review or a compliance audit is handed, and it is worth setting out beside what exists today, because the comparison is where the design either earns its place or does not.

FIGURE 3 · WHAT THE RELIABILITY AUDITOR SEES One question. Two records. Only one of them answers it. TODAY WITH THE GRANT OBJECT Who acted a service account holding supervisory-control access the same service account executed — under a grant issued by a named certified operator; the walk from action to person terminates On what authority an entitlement, plus an approved change ticket a grant identifier that resolves, carrying its issuer, class, apparatus scope and expiry Under which conditions reconstructed afterwards from telemetry and config history the predicate asserted at issuance and the predicate measured at execution, with the gateway's decision recorded between them Was the envelope enforced no — configuration filtered the recommendations; nothing checked the state at the command yes — commands outside the predicate were refused at the gateway, and the refusals sit in the same record as the executions Scoping a defect every command the account issued in the period, every class, every substation every command executed under grants of the affected class — a query, rather than an investigation Revocation revoke the account's supervisory access, severing every automated pathway including the sound ones revoke by handle — one class, or one operator's grants — and leave the rest standing The CIP family governs the perimeter and the access regime — who may reach the system, and how that reach is controlled, monitored and revoked. Both columns comply with it equally. That is the point: the question in the header is a different question. The left column is compliant and answers a different question. The right column is the control and the evidence at once.

Per operation, the record carries: the command and its target; the grant identifier it executed under; the grant's issuer, their certification class and their shift; the conditions asserted at issuance; the conditions measured at execution, with the source of that measurement; and the gateway's decision. Across a period, the record carries the grant history — which grants were open when, issued by whom, closed by clock or by condition, and revoked by whom — plus every refusal.

Two operational consequences follow, and they are the two the teardown identified as the ones that hurt during an event rather than after one.

Revocation becomes surgical. When a study class comes under suspicion, the utility withdraws that class and every grant referring to it closes, everywhere, in one action — while every other automated pathway keeps running. The alternative available today is to revoke the service account's supervisory access, which severs the sound pathways along with the suspect one, in a sector where the automated pathway may itself be load-bearing for reliability. Having something between the blunt outage and doing nothing is not a convenience; it is the difference between a control the utility will actually use during an event and one it will decline to use because using it is worse than the risk.

And population scoping becomes a query. Which operations were taken under the defective authority is answerable by selecting the operations executed under grants of the affected class, in the affected period, which is a database question. Today it is an investigation across every command the account issued, across every class and every substation, and the review must either over-scope at large analytical cost or under-scope on judgement it cannot evidence. In a sector whose entire safety culture rests on events being reconstructed exactly, the difference between a query and an investigation is not an efficiency saving. It is whether the reconstruction is trustworthy.

The bill

Five costs and limits, at full strength, because a construction whose author will not price it is a sales document.

It costs work at shift start, and shift start is not free. Somebody has to issue grants, and the honest version of that is a few minutes of a certified operator's attention at the beginning of every shift, on a job that already has a dense handover. Done badly it becomes a formality with a default template, which is the decay described above. Done well it is a version of something the shift already does — reviewing the day's planned works, the outage schedule and the system conditions — with an output attached. The design should ride on the existing handover rather than adding a ritual beside it, and if an implementation cannot do that, I would expect the grants to become templated within a month.

It adds a check in the action path, and the action path is the thing that had no room. The predicate evaluation happens between proposal and command. It has to be fast — a state lookup and a comparison, not a study — and it has to have a defined behaviour when it cannot complete. I do not have a measured latency for this, because I have not built it inside a production EMS estate and nobody appears to have published one, so anyone quoting a number for it should be asked where the number came from. What I can say is that the check is arithmetic over a state reading the dispatch point already holds for other reasons, and that the sector's own supervisory-control paths already perform checks in this position.

A condition exit can fail unsafe if the condition source is stale, and that is a new failure mode the design introduces. If the gateway's state reading lags reality, a grant can remain open after the system has left the envelope, and the construction will produce a confident record of a permitted action that should not have been permitted — which is worse than no record, because it is a wrong record with an audit trail. The mitigation is a freshness requirement on the reading and fail-closed behaviour when it is not met, and the residual risk is that fail-closed during a degraded telemetry event stops automation exactly when the system is stressed. That trade is real, it is not resolvable in the abstract, and it belongs to the utility's own reliability judgement rather than to an architecture piece.

Where the operator does not own the dispatch point, the design does not reach. If the vendor's platform both produces the recommendation and dispatches it, the independent verifier does not exist and cannot be inserted by the utility. The available responses are contractual — require the dispatch path to terminate in a gateway the utility operates, and require grant verification as a product property — and they are procurement responses rather than engineering ones. I would rather state that plainly than describe a construction that assumes an architecture the buyer may not control.

No instrument requires this, and the fact that instruments exist nearby makes that easy to misstate. FERC's six show-cause orders of 18 June 2026 and its 16 July 2026 direction to NERC — mandatory computational-load reliability standards by 31 December 2026, registry criteria bringing a computational-load entity class into the framework, a Phase II work plan due 1 March 2027 — are the freshest instruments in this sector, and every word of them governs AI as load. The CIP family governs the perimeter and the access regime. The certification programme governs the humans. The November 2024 NERC white paper sets a posture and attaches no compliance obligation to anyone. So the case for building this is not that a standard says so. It is that the standards process takes existing industry practice as its primary input, and that the interval before drafting begins is the only period in which what a utility builds influences what it is later measured against.

And the field is not waiting. ADNOC and SLB have announced an AI real-time operations centre live across ADNOC's full onshore and offshore drilling fleet of more than 120 rigs, hosted inside ADNOC's UAE sovereign cloud, with company-published figures of 30 to 40 percent less engineering effort, engineers covering two to three times more rigs, and incident response times cut by four to twelve hours. Those are the companies' own published numbers about their own deployment and I cite them as such, not as audited measurements. What they establish is the pattern rather than the arithmetic: fleet-scale, centralised, AI-driven operations across physical assets, with sovereignty and governance marketed as part of the release, is what national champions in this sector are shipping now. A construction that arrives after that pattern has propagated is a retrofit.

What would falsify this

Four things, stated at their strongest.

If the independent state source does not exist at the dispatch point. The safety property depends on the gateway holding a view of system conditions that the agent did not produce, fresh enough to be meaningful at command time. If in real utility architectures the consolidated state view is produced by the same platform that hosts the advisory function — which is entirely plausible where the agent is an EMS feature rather than a bolt-on — then the independence is notional and the design needs a different verifier or a different predicate. This is the assumption I am least confident about, and it is the first thing I would test in any real estate.

If a utility already issues machine-readable switching authority and I have not found it. The sector publishes very little architecture. A utility whose gateway already mints per-window execution credentials bound to a certified operator and to stated conditions, and verifies them at the point of work, has the object, and this piece is then a description of what everyone else does. I have found no such deployment described in a primary source, but absence of publication is weak evidence in a sector that publishes as little as this one, and a single public counterexample would confine the argument.

If pre-positioned grants decay as fast as clicks do. The claim that a shift-start decision retains more judgement than a per-action click is a mechanism argument, not a measurement, and I could not find published data on either. If operators issue identical maximal grants every morning within a month of deployment, the construction has bought only visibility of the decay and not any reduction in it — which is still worth something, and considerably less than the piece claims. The test is cheap and would be visible in the grant history within a quarter of any real deployment.

If reviewers in practice accept the change ticket. If post-event reviews and compliance audits treat an approved automation change and a secured account as answering who authorised the operation, then the gap has no institutional consequence and the construction is an engineering preference. I have no basis for claiming reviewers reject it and will not invent one. What I can point at is that this sector already refuses the equivalent answer for people: nobody accepts that the lineman had a key as an answer to who issued the switching order. The argument is that the sector should believe its own doctrine for machines, and that is a normative claim rather than an empirical one.

One limit that is easy to miss. Everything above bounds what the machine may do and records the authority it did it under. None of it makes the recommendation correct. A correctly granted, correctly scoped, condition-verified switching sequence drawn from a study that under-modelled the condition will execute, and the record will show — accurately — that everything was in order. This is a containment and evidence layer, not a quality layer, and the study function's discipline remains exactly as load-bearing as it was. Anyone presenting an authority architecture as a safety guarantee is selling the wrong property.

The smallest useful version is one study class, on one advisory agent, in one control area. Split the credentials so the agent holds nothing that reaches apparatus. Have the shift issue one grant a day. Put the predicate check at the gateway with a fail-closed default. Then run it for a quarter and read the refusals, because the refusals will tell you whether the class was specified correctly — and that is a piece of information the current architecture cannot produce at all.