Research · Output published

Agent governance readiness benchmark

What organisations actually score across entitlements, enforcement, attestation and model risk.

The premise

Every claim about how ready enterprises are for agent governance is currently an assertion. Vendors publish surveys of what people say they intend to do; nobody publishes what a fixed rubric returns when the same questions are put to everyone. This strand exists to replace one anecdote with a distribution.

If you are making a decision

What this changes for a practitioner

  • A median to place your own score against, per sector, rather than a vendor's readiness percentage
  • Which of the four layers is weakest across the field — consistently the ones that require evidence rather than policy
  • A defensible number for a board paper that is not your own self-assessment
For peers · method

How it is actually done

  • A fixed, deterministic rubric: the same answers always produce the same score, which is what makes cross-respondent comparison mean anything at all
  • Runs are stored anonymously — four layer scores, sector, jurisdiction, action class, volume band, and nothing that identifies a respondent, including to me
  • Only runs at 60% completeness or better are counted, so an abandoned assessment cannot drag a sector
  • Medians rather than means: at this sample size one outlier moves an average somewhere indefensible
  • A sector is published only at twelve or more runs; below that the number describes who happened to take the assessment rather than the sector

What is unresolved

Published as unresolved on purpose. A practice that only publishes what it has solved is doing marketing.

Q1

Self-assessment measures self-perception. How far does a scored self-assessment diverge from what an examiner would find — and is the gap itself the more useful measurement?

Q2

Does readiness track regime, organisation size, or how long agents have been in production? The current sample cannot separate these.

Q3

Is there a layer ordering effect — do organisations that start at attestation end up better or worse than those that start at entitlements?

Standards, regimes and dated claims on this page were last checked against their primary sources on . Instruments move — several cited here changed inside the last year — so verify against the source before you rely on one. If you find something stale, tell me and I will correct it.

Working on any of this? Compare notes — in the open or under NDA. All strands: research.

Next step

Where this meets a live decision

Research is upstream of the advisory work. If this strand maps onto something you are deciding now, that is the useful conversation.