Somewhere in your agent programme's future is a Tuesday. An action gets challenged — a customer disputes, a reconciliation breaks, a regulator inquires, or the system simply does something nobody expected.

This is not pessimism; it is arithmetic on non-zero error rates at production volume. A system taking thousands of consequential actions a week will eventually take one that somebody objects to, and the probability of that compounds with every workflow you add.

So the executive question is not whether that Tuesday arrives. It is which of two programmes it arrives at — and that is decided now, in the boring quarter before, not in the war room during.

The same Tuesday, two programmes The incident is identical. What differs was decided in the boring quarter before. With the artefacts T+0 challenged +1 min authority register: was this even possible? +1 hr attestation record: review scoped to ONE action same day matrix = lookup, not debate kill path: one agent, not all narrow, evidenced disclosure The programme keeps running throughout. The incident becomes a footnote. Without them T+0 challenged days 1–3 was it permitted? → a debate about intent week 1 can't bound what else it did review expands to everything week 2+ only lever: turn it off the incident IS an outage Disclosure is of an unknown, not an incident. Every conversation is longer. The freeze is the rational response, so the freeze happens. Programmes rarely die of the wound. They die of the diagnosis taking too long.

I have watched both versions from the inside, and the gap between them is not competence or seriousness. Both programmes were run by capable people. The difference is five artefacts and one rehearsal, all of which are cheap, and none of which can be produced under time pressure.

Why the first hour decides the shape of everything after it

Before the artefacts, the mechanism — because it explains why small documents have outsized effects.

An incident has a scope, and the scope is set very early by what can be established quickly. If you can show, in the first hour, that a specific action was within a specific grant and that you can see exactly what informed it, the incident stays an incident: one action, one review, one remediation. The programme keeps running.

If you cannot, something different happens, and it is entirely rational. Nobody can bound what else the system might have done. In the absence of a bound, the responsible assumption is the worst one, so the review expands from the challenged action to the action class, then to the agent, then to everything the agent could reach. Each expansion is defensible. Collectively they produce a months-long investigation and, usually, a freeze — because freezing is the only control available to someone who cannot scope the exposure.

That is why the line worth carrying to a board is that programmes rarely die of the wound. They die of the diagnosis taking too long. The initial error is often small; the inability to characterise it is what becomes existential.

Artefact one: the authority register

One page per agent action class: what it may do, its ceilings, the named grantor, the expiry.

The five artefacts Each one converts a question that takes days into a question that takes minutes. 1 The authority register One page per action class: what it may do · ceilings · named grantor · expiry. Answers: “was this action even supposed to be possible?” — in one minute. 2 The attestation record What it knew, from where, under which entitlement, past which check — at act time. Keeps the review scoped to ONE action instead of expanding to everything. 3 The kill path — rehearsed Revoke one agent's authority in minutes, mid-run, without downing the workflow. Never executed under pressure? Then it is an intention, not a control. 4 The decision matrix — pre-agreed Which classes pause the action, which pause the agent, which pause nothing. Decided in daylight. Converts the worst meeting of the quarter into a lookup. 5 The notification map Who is told, when, by whom — internally, and externally where obligations attach. Drawn with counsel now, because the clock does not wait for counsel later. Artefacts one and two do most of the work: they make every downstream conversation shorter and every disclosure narrower. Evidence is the difference between reporting an incident and reporting an unknown.

In the incident, this is the document that answers the first executive question — was this action even supposed to be possible? — in one minute instead of one taskforce. It is a binary that immediately partitions the problem: either the system did something it was granted and the grant was wrong, or it did something it was not granted and the enforcement failed. Those are completely different investigations with different owners, and knowing which one you are in is worth more than any other single fact.

Programmes without it discover, mid-incident, that they are litigating their own intent. The conversation becomes archaeological — what did we mean when we set this up, who approved it, was this in scope — conducted by people reconstructing decisions from memory and Slack threads while the clock runs.

The register is also the artefact most often assumed to exist. Someone will say the permissions are in the platform. Permissions in a platform are a technical configuration, not a statement of authority: they say what the system can do, and the register says what it may do and who decided that. The gap between those two is where the incident lives.

Artefact two: the attestation record

What the system knew, from where, under which entitlement, past which check — captured at act time, mechanically.

Two words in that sentence carry the weight. At act time, because a record assembled afterwards from logs is a reconstruction, and a reconstruction is exactly what is in dispute. And mechanically, because anything depending on a person remembering to record it will be absent for the action that matters.

This is the artefact that keeps the review scoped. With it, you can say: this action, this context, this grant, this check — and everything outside that boundary is demonstrably unaffected. Without it, the boundary cannot be drawn, and the expansion described above follows automatically.

The common substitute is logging, and the distinction matters more here than anywhere else in the stack. Logs are produced for debugging and retained on a debugging schedule — often thirty or ninety days, which is shorter than the interval between an action and a challenge to it. Attestation is produced for reconstruction and retained on an investigation schedule. A programme that discovers this distinction during an incident discovers it by finding that the relevant records aged out.

Artefact three: the kill path — rehearsed

Can you revoke one agent's authority in minutes, mid-run, without taking down the workflow for everyone?

If the only lever is turn the platform off, then every incident is automatically an outage, and your incident cost silently includes your uptime. That changes the economics of responding: a team facing an all-or-nothing switch will hesitate, escalate, and seek certainty before acting — which is precisely the wrong incentive in the first hour.

Granularity is therefore a safety property, not a convenience. The ability to revoke one action class for one agent means the response can be proportionate, which means it can be fast, which means the exposure window closes.

And the rehearsal clause is not decoration. If revocation has never been executed under pressure, you do not have a control; you have an intention. Revocation paths rot the way failover paths rot — a credential expires, a dependency changes, the runbook references a console that was replaced. Rehearse it the way you rehearse failover, because that is what it is.

Artefact four: the decision matrix, pre-agreed

Which incident classes pause the action class, which pause the agent, which pause nothing — decided in daylight, with risk, legal and the business in the room.

The value here is entirely about when the decision is made rather than what it is. In-incident escalation debates consume exactly the hours that determine whether the incident is contained or discussed publicly, and they are conducted under conditions engineered for bad judgement: incomplete information, visible stakes, and participants whose functions have genuinely different interests.

Pre-agreeing it converts the worst meeting of the quarter into a lookup. It also changes who is exposed: with a matrix, the person deciding to pause is executing a policy the organisation agreed to; without one, they are making a career-shaped judgement call at speed. That asymmetry explains a great deal of the hesitation you see in unprepared programmes.

One design note. The matrix should be keyed to incident class rather than to severity estimate, because severity is exactly what you cannot assess in hour one. Classes are observable — an action outside grant, an action inside grant with a disputed outcome, a data-exposure event, an availability event — and each maps to a pre-agreed posture. Severity-keyed matrices push the hard judgement back into the moment they were meant to remove it from.

Artefact five: the notification map, and the clocks nobody has read

Who is told, when, by whom: internally, and externally where obligations attach.

In regulated sectors those obligations attach fast, and the crucial property is when the clock starts. It starts at awareness, not at understanding — which is a very uncomfortable asymmetry, because at the moment of awareness a programme typically has an alert and a hypothesis, while the obligation wants scope, cause and remediation.

The clocks start at awareness, not understanding India as the worked example. Most regulated sectors have a version of this. T+0 · you become aware you understand almost nothing yet CERT-In · 6 hours Directions of April 2022. Certain cyber-security incidents reportable within six hours of detection. Six hours. DPDP Rules 2025 · 72 hours Board informed without delay; then within 72 hours: nature and extent, circumstances, remedial measures, findings on who caused it. Affected Data Principals: 72 hours. One breach, two independent clocks. And the 72 hours runs continuously weekends, public holidays, outside business hours. What the clock demands Scope — what was affected Cause — how it happened Remediation — what you did What an unprepared programme has An alert A hypothesis And a taskforce being assembled The notification map is drawn in advance because the clock starts before anyone understands anything.

India is the sharpest worked example, and it is instructive because one event can start two independent clocks. Under the CERT-In Directions of April 2022, certain cyber-security incidents are reportable within six hours of detection. Under the DPDP Rules 2025, the Data Protection Board must be informed without delay, with a fuller account — the nature and extent of the breach, the circumstances, the remedial measures taken, and findings as to who caused it — inside seventy-two hours, and affected Data Principals notified on the same seventy-two-hour horizon. That window runs continuously: weekends, public holidays, outside business hours.

Six hours is the number that should reorganise a programme's thinking. It is shorter than most incident bridges take to establish what happened. It is far shorter than the time required to assemble a taskforce, and it does not care that your understanding is partial.

Which is the entire argument for drawing the map in advance with counsel. The clock does not wait for counsel later, and the decisions the map encodes — what constitutes awareness for our purposes, who is authorised to file, what a first filing says when facts are incomplete — are decisions nobody should be making for the first time at hour four.

And this is where artefacts one and two pay for themselves at the highest rate. A programme that can produce them makes every one of those conversations shorter and every disclosure narrower, because evidence is the difference between reporting an incident and reporting an unknown. A regulator receiving a bounded account of one action is in a different conversation from a regulator receiving a statement that something happened and the scope is under investigation.

The rehearsal, and what it actually tests

One tabletop, ninety minutes, twice a year. Inject a plausible scenario and walk the five artefacts in anger.

Ninety minutes, twice a year The rehearsal is what stitches the five artefacts into a capability. One tabletop · 90 minutes · twice a year Inject a plausible scenario and walk the five artefacts in anger: “the agent honoured a policy that had been superseded” “the agent's context contained something it should never have seen” 1 Register Can you produce it for the affected action class — without asking anyone? 2 Attestation Retrieve the record for a specific action from three months ago. Time it. 3 Kill path Revoke one agent's authority now, clock running, without downing the workflow. 4 Matrix Does it resolve this case — or does the room start arguing? The arguing is the finding. 5 Notification Could the map actually be executed inside six hours? The first run is always humbling. That is the point. Cheaper humble in a conference room than in front of an examiner. Credit risk's playbook. Cyber's playbook. A new actor. Nothing novel — so nothing excusable to skip.

Two injections I would use, because both are realistic and neither is exotic. The agent honoured a policy that had been superseded — which tests whether anyone can establish what the system believed at the time. And the agent's context contained something it should never have seen — which tests the boundary between the context layer and the entitlement model.

What the tabletop measures is not whether the artefacts exist but whether they can be used under time pressure by the people who would actually be in the room. Those are different properties. A register nobody can locate, an attestation store nobody has queried, a revocation runbook referencing a decommissioned console — all of these exist on paper and fail in practice, and the only way to discover that cheaply is to try.

The first run is always humbling. That is the point: cheaper humble in a conference room than in front of an examiner. And the specific finding to watch for is not a missing document — it is the moment the room starts arguing about who decides. That argument is the artefact-four gap, surfacing exactly where it would have surfaced during the real thing.

Who owns each artefact — the question that decides whether they stay current

Every artefact here has a natural owner, and getting the assignment wrong is the most common reason they exist on paper and fail in practice.

The authority register belongs to the business owner of the action class — the person who would be accountable for the outcome if a human had done it. Not to engineering, which knows what the system can do but not what it should be permitted to do, and not to risk, which can challenge a grant but should not be authoring it. The register's whole value is that it records a business decision, so a register authored by a technical function is recording the wrong thing.

The attestation record belongs to engineering, because it is a system property. The useful constraint is that its retention schedule should be set by legal rather than by whoever configures the log pipeline — this is the single change that prevents the most common failure, which is records ageing out before a challenge arrives.

The kill path belongs to whoever runs production. The decision matrix belongs jointly to risk and the business, and jointly is doing real work in that sentence: a matrix authored by risk alone gets ignored under pressure, and one authored by the business alone will not survive contact with a regulator. The notification map belongs to legal, informed by everyone else.

The diagnostic is simple: if you cannot name a person for each of the five, they will drift. Documents with a function attached rather than a person are documents nobody updates, because updating them is nobody's specific job.

What week two actually looks like

The claim that this is week-two work rather than a quarter's project deserves substantiation, because it sounds like consultant optimism.

For a programme with two or three agent action classes — which is what most programmes actually have, whatever the roadmap says — the register is three pages. Each page answers four questions: what may this do, up to what limit, granted by whom, expiring when. The hard part is not writing it; it is that answering the questions forces decisions that were previously implicit, and those conversations take a couple of hours each.

The attestation record is the only item with genuine engineering content, and even that is smaller than it sounds at the pilot stage: capture, at the moment of action, the identifiers of what was retrieved, the grant relied on, the check result, and a timestamp. That is a structured log line with a defined retention. It is not a platform.

The kill path usually already exists in some form and has simply never been exercised, so the work is a rehearsal rather than a build. The matrix is one workshop. The notification map is one session with counsel plus a page.

So the honest estimate is days of work spread across two or three weeks, most of it conversation rather than construction. What makes programmes skip it is not the cost — it is that none of it produces anything demonstrable, while the same fortnight spent on capability produces a demo. Naming that trade-off out loud is most of the battle, and it is the same trade-off the sourcing argument runs into.

Three markets, three different failure points

The five artefacts are constant. Which one is missing when I arrive is not.

In North America, the artefacts most often exist in some form, because the incident-response muscle is mature and cyber has already forced the pattern. The failure point is artefact three: the kill path assumes a platform-level switch, because that is what the vendor provides, so revocation is all-or-nothing and every incident threatens uptime. The fix is usually a conversation with the platform vendor about granularity that nobody has had.

In India, the binding constraint is the clock rather than the artefacts. Six hours to CERT-In and seventy-two to the Board is an operationally demanding combination, and the failure point is artefact five — a notification map that exists as a legal memo rather than as an executable runbook with named people and out-of-hours contacts. The rehearsal question to ask here is not do we have a map but could we file inside six hours on a Saturday.

In the Gulf, new-build programmes tend to have the cleanest architecture and the thinnest incident history, which produces a specific gap: artefact four. There is no accumulated organisational memory of what a pause costs, so nobody has pre-agreed when to pause, and the first incident becomes the occasion for inventing the escalation policy. That is the most expensive possible moment to invent it, and it is entirely avoidable in a single workshop.

The strongest objection: this is theatre for an incident that may never come

The serious counter-argument deserves stating, because it is what a sceptical executive is actually thinking.

It runs: this is a lot of process for a hypothetical. Most agent deployments are low-stakes — drafting, summarising, internal retrieval — and will never generate an incident anyone outside the team hears about. Building incident machinery before the system has proved value is exactly the governance-first sequencing that makes AI programmes slow and expensive, and the five artefacts will become stale documents nobody maintains, which is worse than not having them because they create false assurance.

The staleness point is correct and is the real risk. Documents that describe a system as it was eighteen months ago are actively harmful in an incident, because they will be trusted. Any programme adopting this should treat the artefacts as living or not adopt them — which is, incidentally, the strongest argument for the twice-yearly rehearsal, since a rehearsal is the cheapest mechanism anyone has found for detecting that a document has drifted from the system.

The low-stakes point is where it breaks, in a specific way. The artefacts are proportionate to what the agent may do, not to what the programme costs. An agent that only drafts and summarises genuinely needs a very thin version — a one-line register saying it may read these sources and write nothing, and that is nearly the whole exercise. The cost is small precisely because the authority is small. What is not defensible is skipping the exercise entirely, because then nobody has established that the authority is small, which is the very claim the sceptical position rests on.

There is also a scope-creep dynamic that makes the low-stakes assumption unstable: systems that begin as drafting tools acquire actions. Nobody re-runs the risk assessment when a summariser gains the ability to send. The register is what makes that acquisition visible.

Where this argument is weakest

Two admissions.

First, the five artefacts are drawn from practice and from adjacent disciplines rather than from a study showing that programmes possessing them experience better incident outcomes. That study does not exist for agentic AI — the deployment base is too young and incidents are not systematically reported. The confidence here comes from the mechanism being legible and from the structure having worked in credit risk and in cyber, not from outcome data in this domain.

Second, the regulatory detail above is India-specific and current as described; the general claim that clocks start at awareness holds broadly, but the specific hours do not transfer. Anyone applying this outside India should have counsel identify their own clocks rather than assuming six and seventy-two.

Two questions for your next review

Board members reading this will have seen the structure before — it is credit risk's playbook and cyber's playbook, applied to a new actor. That is both the reassurance and the indictment: nothing here is novel, so nothing here is excusable to skip.

The programmes that treat the first incident as a when build these five artefacts in week two. The programmes that treat it as an if meet me in a very different meeting.

So: ask for the five artefacts at your next review. Then ask when the last rehearsal was. Those two questions cost nothing today and roughly everything on that Tuesday.