Forward-deployed architecture, governance and assurance

Don’t take my word for it. Make the policy refuse something.

Your AI agents will be audited. I make sure they survive it.

Every enterprise AI program eventually meets one question: which agent did that, on whose authority, and can you prove it? Below is a working evaluator answering it — not a recording, not a mock-up.

Policy evaluator · running heredeterministic

Change either field below. The decision, the evidence record and its digest all change with it.

DENY · R2 · attributable authority

The grant covers the action, but it is a shared account — so no reviewer could later say which person authorized this. Consequential actions require an authority that resolves to someone.

Evidence record — what a reviewer would find months later

action        adjust_limit
requires      write:limits
principal     broad
attributable  no — shared account
checkpoint    absent
rule          R2 · attributable authority
decision      DENY

digest a5a8df4b · changes with any field above. Tamper-evidence demonstrated, not a signature — there is no key in your browser.

The claim, and the thing behind it

Refusal is the easy half

Any gateway can block a call. What almost no stack produces is the record that makes a refusal — or a permission — defensible to someone asking about it eleven months later.

That is the difference between a control and evidence of a control, and it is the difference an examiner is trained to find.

What a reviewer asks for
which agent      → identity, not "the platform"
whose authority  → resolves to a person
what it touched  → data, tools, downstream
what stopped it  → the rule, by name
when             → ordered, immutable
who reviewed     → checkpoint, or absence of one
The method

Four layers, in dependency order

Each depends on the one above it. The evaluator above only works because the first two are in place — the third is what makes it survivable afterward.

01

Entitlements

An agent can only act inside an authority someone actually granted it.

02

Policy enforcement

At the point of action — not in a document nobody reads.

03

Attestation & audit

Months later, reconstruct what happened and why it was permitted.

04

Model risk governance

That satisfies a reviewer, not just a dashboard.

Evidence, not testimonials

No client names. So the work has to speak.

There are no logos on this site and there never will be. What exists instead is a research program anyone can inspect — a readiness benchmark built from anonymous assessments, open-weight parity evaluation, incident patterns read from primary disclosures, and a published list of the problems nobody has solved.