Diagnostic

Would your agents survive an audit?

Four modules — readiness scoring, obligation mapping, evidence gaps and modeled exposure. It starts at eight questions and goes deeper only where your answers reveal a gap.

No email required. Nothing you enter identifies you or your organization. An anonymous copy of the four scores is retained to build a sector benchmark — no organization name, no identifier, no IP address, no cookies.

This site tracks nothing that identifies you — no cookies, no session recording, no person profiles. The same standard it assesses you against, applied to itself: here is the evidence.

Question 1 of 8

Context

Which regime governs you most directly?

Method, in full

What the diagnostic tests

The interactive assessment above is backed by a deterministic scoring model. The same answers always produce the same score. These are the questions and decision rules in prose so a reviewer—or a crawler that does not execute JavaScript—can examine the method.

Entitlements

Core test: When an agent takes an action, can you name the human authority it derives from?

Most agent stacks inherit a service account holding the union of everyone's permissions. Authority has to be granted, scoped, time-boxed and recorded — per agent, per action class, traceable to the human who delegated it.

Follow-up depth: What do your agents' credentials actually permit? · Does an agent's authority expire or get re-examined? · Are business limits (value ceilings, rate caps, jurisdiction) enforced on the action, or assumed in the prompt?

Policy enforcement

Core test: Where is policy evaluated for an agent action?

A control that runs at login and never again is not a control. Enforcement belongs at the moment of the action, between the agent's intent and the system of record, where it can still refuse.

Follow-up depth: Can the enforcement layer actually refuse an action in flight? · When an agent makes a customer-facing commitment, is it checked against an authoritative source? · For third-party agent platforms, can you enforce your own policy inside their execution path?

Attestation & audit

Core test: Could you reconstruct, six months from now, exactly why a specific agent action was permitted?

Chat transcripts are not traceability. Evidence means the decision, the authority it rested on, the policy evaluated, and the inputs — retained and reconstructible on a reviewer's timeline, not yours.

Follow-up depth: Do your records capture which model and prompt versions were in effect at the moment of the action? · Is evidence retained for as long as your regime requires? · Where a human is 'in the loop', can you evidence what they actually saw and how long they had?

Model risk governance

Core test: How do your agents appear in model risk inventory?

Agent behavior changes the day a provider ships a new model version, with no code change on your side. Inventory, validation, effective challenge and monitoring all have to account for a system that acts rather than predicts.

Follow-up depth: When a provider ships a new model version, what happens? · Can an independent reviewer mount effective challenge to an agent's behavior? · If an agent starts acting wrongly at scale, what stops it?

BandScoreInterpretation
Critical gap0–29Authority or evidence is absent; a review is likely to stop here.
Weak30–54Controls exist informally but cannot be reconstructed or enforced consistently.
Developing55–79Core controls exist; gaps remain in coverage, evidence or independent challenge.
Defensible80–100The design is likely defensible, subject to testing against a real historical action.

Context that changes the obligation map

  • Which regime governs you most directly?
  • Where does the obligation land?
  • What is the most consequential thing an agent can do today without a human approving it?
  • Roughly how many agent-initiated actions occur per month?

Entitlements and attestation receive additional weight because enforcement and model-risk controls depend on knowing who authorized an action and being able to reconstruct it later. The exposure range is directional prioritization—not a penalty forecast.