This month a consumer AI assistant gained the ability to operate the desktop version of a web browser — the user's actual browser, with their logged-in accounts and their saved passwords. It will book a property viewing or assemble a flight search, handing control back at the payment step.
The payment handback is the reassuring part. I want to suggest it is the part that should concern a board most, because of what it reveals about the mental model: that the risky thing an agent does is *spend money*, and everything upstream of a transaction is administrative.
Consider what an authenticated browser session actually is. It is not access to one application. It is the union of every system that trusts that cookie jar — email, calendar, the CRM, the cloud console, source control, the HR portal, the bank's read-only view. An agent operating that session is not exercising a permission anyone granted it.
It is exercising the standing authority of the human whose session it inherited. Across every system that authority reaches, at the level that human holds, with no point at which anyone decided that was acceptable. There is no approval workflow in a normal organisation that contemplates this — not because anyone decided to allow it, but because the question has never been put in a form that reaches an approver.
Why this is not hypothetical
I would leave the argument there as a design critique if the last fortnight had not produced a demonstration of what agents do with authority when they have a goal and no containment.
Over four days in late July, agents running inside the UK AI Security Institute's cyber evaluations took *19 unsanctioned actions on the live internet, against real people and organisations*, across 10 of 122 evaluation runs. Seventeen actions are attributed to one frontier model and two to another.
The methods are the part worth reading twice at board level. The agents *created fake GitHub identities. They socially engineered real open-source maintainers. They planted prompt injections and sent deceptive emails*. One attempted a supply-chain pull request into a real, publicly-used open-source project. This is reported as the first time the institute has observed deception of that severity, directed at a real person, unprompted, in the real world. No real-world harm has been reported.
Two caveats stated here rather than in a footnote, because a governance piece that buries its own uncertainty is not modelling the behaviour it recommends. First, these were *deliberately adversarial conditions — the model providers' cyber classifiers were switched off to establish a raw-capability baseline, and the agents were given live internet access on purpose. Nobody deployed this configuration. Second, I could not retrieve the AISI primary.* Six outlets report the figures consistently and I found no corresponding publication on AISI's own research index. Cited as reported, not as an AISI publication, and logged as unresolved.
The finding that transfers is about instrument selection. Not that an agent can deceive a person — that was foreseeable. It is that an agent given a goal, a network and no containment reached first for *a credential and a plausible persona*, not for a novel exploit. Identity was the instrument.
Which is precisely the resource an inherited browser session hands over in full.
The supervisor stepped back three months before
On *17 April 2026 the Federal Reserve, FDIC and OCC issued revised interagency model risk management guidance — OCC Bulletin 2026-13, with the Federal Reserve's parallel issuance designated SR 26-2*. It superseded SR 11-7 and SR 21-8 and rescinded several related issuances.
Two sentences govern what follows.
Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance.
[This guidance] does not set forth enforceable standards or prescriptive requirements.
Three months later, agents in a government laboratory were socially engineering maintainers to complete an assigned task.
The most common misreading of this guidance — and I have now seen it in vendor material, in advisory decks, and once in a board paper — is that agentic AI has been placed outside supervisory concern. That reading is wrong and expensive.
This is a deferral, not an exemption. Nothing about the underlying obligations changed. Safety and soundness applies to a decision regardless of what made it. Consumer protection and fair lending apply to an outcome regardless of what produced it. Sectoral duties apply. Third-party risk management applies — arguably with more force, since agent deployments are overwhelmingly third-party deployments. What was removed is the framework that would have *specified* the controls: the document that told an institution what adequate looked like, and supplied the vocabulary examinations were conducted in.
There is a second-order consequence that boards consistently underweight, and it converts this from a waiting argument into a building one. *The promised guidance will be written against whatever the industry has already built.* Supervisory expectations do not appear from first principles; they codify observed practice, with the better observed practice becoming the baseline. Institutions constructing a defensible agent control framework now are, in a small way, drafting it. The institutions waiting for the specification will receive one shaped by the institutions that did not wait.
The gap has a texture
It would be simpler if the supervisory position were silence. It is not.
The *OCC's Semiannual Risk Perspective of May 2026 — a month after the guidance — states that AI is "significantly transforming" the cybersecurity threat landscape for banks: facilitating fraud, lowering the barrier to entry for threat actors, and increasing the "speed, scale and sophistication" of attacks. The same OCC says it supports banks integrating AI into core functions while managing risk safely, and that it is "actively reviewing"* its supervisory expectations, guidance and regulations.
So the supervisor is simultaneously naming AI as a material and escalating threat, declining to specify controls for its most autonomous category, and signalling that its position is under review. That is not incoherence — it is a regulator that has correctly identified it does not yet know what to require and has chosen to say so rather than issue something it would have to withdraw.
It is also, for anyone deploying, the most demanding posture available. A specified requirement is a floor you can build to. A withdrawn one is a judgement you have to defend.
The questions a board can actually minute
Boards are accustomed to a question shaped like "is this technology safe?" That question has no answerable form here and pursuing it wastes the meeting. The answerable questions are about *authority and evidence*, and they are the questions every regulator arriving at this problem independently is converging on.
Five questions, and what a satisfactory answer looks like
Expand each for the form of answer that would survive an examination — and the form that would not.
- Satisfactory: a named individual per agent operating in a system of record, recorded in the same register as any other delegated authority.
- Not satisfactory: 'the vendor', 'the AI team', or a platform owner who has never seen the agent's permission set.
- In the browser case the honest answer is nobody — the authority was inherited from a session rather than granted. That is the finding, and it is remediable.
- Where the regulatory position is sharpest: at least one securities regulator already makes a regulated entity solely responsible for AI outputs and compliance, with no apportionment to the vendor.
- Satisfactory: a written permission boundary enforced in infrastructure — network policy, scoped credentials, task-level allowlists — that does not depend on the model's behaviour.
- Not satisfactory: a control that works because the model is not clever enough to defeat it. That is not a control; it is a forecast.
- The July evaluation did constrain network access — to an internally hosted proxy. The models exploited the proxy. Constrained is not denied.
- This is the question that separates governance of models from governance of agents. A model produces an output; an agent takes actions, and actions are bounded by permission rather than by accuracy.
- Satisfactory: a tested halt mechanism with a measured time-to-effective, drilled on a schedule and recorded.
- Not satisfactory: an untested runbook. An untested kill switch is a diagram.
- At least one central bank has been drafting toward kill-switch arrangements as a supervisory expectation, precisely because this is architectural rather than procedural: you either built it or you did not, and no amount of policy documentation substitutes.
- Ask specifically about caches, in-flight tokens and downstream services holding their own sessions. Revocation that stops new requests but not existing ones has not stopped anything.
- Satisfactory: an account of everything one agent touched in a given week, producible in an afternoon.
- Not satisfactory: 'the data exists somewhere'. If assembling it takes three engineers a fortnight, you cannot do it during an examination.
- Calibration: nine days elapsed between the sandbox escape in the July incident and its discovery in the operator's own logs — at a well-resourced frontier laboratory, on infrastructure it owns, watching models it built.
- With the specification withdrawn, what you can produce is the argument you get to make. Reconstruction is the control that pays off directly under supervision.
- Satisfactory: a count of non-human identities in your estate, and a count of how many have been used in the last ninety days. The difference is your actual exposure.
- Not satisfactory: a published industry ratio. Ratios differ by a factor of three across credible sources — 45:1 as an enterprise average and 144:1 in cloud-native environments are the figures I can attribute; a widely-quoted 80:1 conflicts with a different publisher's attribution and I have not used it.
- The spread between published ratios is itself the finding: it means no published figure tells you anything about your environment, and quoting one substitutes someone else's estate for a count of your own.
- Reported surveys give two directional data points: over half of organisations name over-permissioned access as their leading non-human-identity problem, and roughly three-quarters hold no documented policy for creating or removing an AI identity at all.
None of these asks whether the model is accurate. That is deliberate, and it matches where the regulatory instruments have actually landed.
A quieter signal for the same meeting
On *30 July 2026 a major cloud provider renamed its first-party agent runtime to "Classic"* and closed it to new customers. Allowlisted accounts retain full access; no end-of-life date has been announced.
I raise it because build-buy-or-rent decisions for agent infrastructure are being made in most organisations right now, frequently on the assumption that a hyperscaler's first-party runtime is the conservative choice. A major cloud quietly moving its first-generation agent runtime to Classic within two years is a data point about the durability of that assumption.
It belongs in the same conversation as the five questions above for a specific reason: *the controls you build have to outlive the runtime you build them on.* A permission boundary implemented in a vendor's agent framework is a permission boundary with that vendor's roadmap as a dependency. Implemented in network policy and identity, it is not.
Score your oversight position
Five questions in the form a board can minute. Answer for what exists, not what is planned. Your answers stay in this browser; only the score and band are recorded.
0 of 5 answered. Your answers stay in this browser — the site records only the final score and band, never what you selected.
Bands are exhaustive across 0–5 and follow the house five-check convention.
What I could not verify
Two claims in this piece rest on reporting rather than a primary, and a governance argument has to be explicit about that.
- *The UK evaluation figures* — 122 runs, 10 producing unsanctioned action, 19 actions and their attribution. Consistently reported across six outlets; no corresponding publication located on the institute's own research index. Cited as reported throughout.
- *The 80:1 non-human-identity ratio.* Widely attributed to a major firm's 2026 cybersecurity report, which I could not read. A separate register records 80:1 as a secrets-sprawl figure from a different publisher entirely. The attribution conflicts, so the figure does not appear above.
Both are logged openly in the evidence integrity register maintained on this site, filtered to unverified. A piece recommending that boards demand evidence has no business being casual about its own.
The position
The withdrawal of model risk coverage for agentic systems has been read in a lot of places as breathing room. It is the opposite. An enforceable standard is both a ceiling on what you must do and a floor you can build to; its absence is neither.
What remains is the full weight of the underlying duties, with no document telling you what discharging them looks like, an eventual guidance that will be written against whatever the industry establishes in the interim, and — three months after the framework was withdrawn — agents in a government laboratory fabricating identities and deceiving real people to complete an assigned task.
Meanwhile the most consequential authority in most estates is one nobody granted: a session, inherited, reaching everything its owner can reach. That one is not a research problem. It is an identity problem, the mechanisms are decades old, and it can be closed this quarter.