AI Agent Governance & Assurance

AI that can be handed real work — and prove what it did.

I build and govern the whole agentic stack inside regulated enterprises: agent protocols, permission-aware context, security of the stack itself, sovereign deployment, and the governance layer that makes any of it defensible. Your AI agents will be audited. I make sure they survive it.

Adaptive, four modules, about seven minutes. No email required.

Research

The work behind the advice

No client names, no case studies — so the evidence has to be the work itself. A readiness benchmark built from anonymous assessments, open-weight parity evaluation, incident patterns read from primary disclosures, and a published list of the problems nobody has solved.

Output published

Agent governance readiness benchmark

What organisations actually score across entitlements, enforcement, attestation and model risk.

Running

Open-weight parity evaluation

Whether a model you control can match a hosted one on the specific workflow, measured rather than argued.

Running

Agentic incident patterns

What actually went wrong in public agent incidents, reduced to the failure that carried it.

Open question

Open problems in agent governance

The questions this practice does not have good answers to, published deliberately.

The problem

Most enterprise AI programmes don’t fail on model quality.

They fail the first time someone asks: which agent did that, on whose authority, and can you prove it?

That question is arriving now. Agents have moved from drafting text to taking actions — moving money, adjusting limits, touching patient records, filing claims. The moment an agent acts, it inherits every obligation that governs the human it replaced. Most stacks were never built to answer for that.

The stack

Five surfaces. One question underneath all of them.

Agent autonomy, governance and the protocols underneath are what make advanced agent work in an enterprise possible at all. These are the surfaces that work touches — and on every one of them the question a reviewer arrives with is the same: which agent did that, on whose authority, and can you prove it?

Agent protocols & orchestration

The plumbing that decides whether multi-agent work is reliable or merely impressive.

Context graphs & permission-aware retrieval

Grounding agents in enterprise knowledge without quietly dismantling a decade of access control.

Security of the agentic stack

Trust boundaries, egress, secrets and adversarial pressure across the whole agentic and GenAI landscape.

Sovereign & open-weight AI

Owning the intelligence and the context, rather than renting both and hoping.

Model risk & validation for agents

Validation frameworks written for systems that score, applied to systems that act.

The method

Four layers, in dependency order

Each depends on the one above it. Programmes that start at the bottom rebuild everything.

01

Entitlements

An agent can only act inside an authority someone actually granted it.

Most agent stacks inherit a service account holding the union of everyone's permissions. Authority has to be granted, scoped, time-boxed and recorded — per agent, per action class, traceable to the human who delegated it.

02

Policy enforcement

At the point of action — not in a document nobody reads.

A control that runs at login and never again is not a control. Enforcement belongs at the moment of the action, between the agent's intent and the system of record, where it can still refuse.

03

Attestation & audit

Months later, reconstruct what happened and why it was permitted.

Chat transcripts are not traceability. Evidence means the decision, the authority it rested on, the policy evaluated, and the inputs — retained and reconstructible on a reviewer's timeline, not yours.

04

Model risk governance

That satisfies a reviewer, not just a dashboard.

Agent behaviour changes the day a provider ships a new model version, with no code change on your side. Inventory, validation, effective challenge and monitoring all have to account for a system that acts rather than predicts.

Built to NIST AI RMF · Federal Reserve SR 11-7 · HIPAA · India DPDP
and the EU AI Actwhere a client’s footprint reaches Europe.
Where I’m useful

Banks, insurers, health systems and sovereign platforms standing up agents in environments where a wrong action is a reportable event. Usually I am called in when a programme is moving fast and someone senior has realised the governance layer was assumed rather than designed.

Where I’m not

If you need a prototype to demo next month, I am the wrong call. My work only pays off when the system has to survive scrutiny.

Next step

If any of this is live for you

The fastest route is a conversation. Bring the architecture you are worried about — the first useful thing usually surfaces inside twenty minutes.

Not ready for a call?

Send one policy, model card or vendor claim. I will return a written governance critique within three business days—one free teardown per organisation.

Request an asynchronous teardown