Entitlements
An agent can only act inside an authority someone actually granted it.
Your AI agents will be audited. I make sure they survive it.
Every enterprise AI program eventually meets one question: which agent did that, on whose authority, and can you prove it? Below is a working evaluator answering it — not a recording, not a mock-up.
Change either field below. The decision, the evidence record and its digest all change with it.
The grant covers the action, but it is a shared account — so no reviewer could later say which person authorized this. Consequential actions require an authority that resolves to someone.
Evidence record — what a reviewer would find months later
action adjust_limit requires write:limits principal broad attributable no — shared account checkpoint absent rule R2 · attributable authority decision DENY
digest a5a8df4b · changes with any field above. Tamper-evidence demonstrated, not a signature — there is no key in your browser.
Any gateway can block a call. What almost no stack produces is the record that makes a refusal — or a permission — defensible to someone asking about it eleven months later.
That is the difference between a control and evidence of a control, and it is the difference an examiner is trained to find.
which agent → identity, not "the platform" whose authority → resolves to a person what it touched → data, tools, downstream what stopped it → the rule, by name when → ordered, immutable who reviewed → checkpoint, or absence of one
Each depends on the one above it. The evaluator above only works because the first two are in place — the third is what makes it survivable afterward.
An agent can only act inside an authority someone actually granted it.
At the point of action — not in a document nobody reads.
Months later, reconstruct what happened and why it was permitted.
That satisfies a reviewer, not just a dashboard.
There are no logos on this site and there never will be. What exists instead is a research program anyone can inspect — a readiness benchmark built from anonymous assessments, open-weight parity evaluation, incident patterns read from primary disclosures, and a published list of the problems nobody has solved.