On the fifth of May 2026, at a bank in Pennsylvania, somebody had a task to get through and used a tool that made it easier. The tool was an AI application the bank had not approved. What went into it was non-public customer information: names, Social Security numbers, dates of birth. There is no evidence anywhere in the public record that this person was doing anything other than trying to finish their work.
Two days later, on the seventh, CB Financial Services determined that the incident was material — not because operations were disrupted, because they were not, and not because it was expected to affect the financial results, because the company said it was not, but because of what the filing calls the volume and sensitive nature of the non-public information at issue. That determination started a four-business-day clock. On the eleventh, the Form 8-K was on EDGAR under Item 1.05, and a routine afternoon at a community bank became the first materiality determination in the United States triggered by an employee's unauthorised use of AI rather than by an attacker.
I open with that story and then have to take it away, because the honest thing about it is what it is not. Nothing in that incident acted. No system took a decision, initiated a payment, wrote to a record or committed the bank to anything. It is a shadow AI incident — a person, a tool, and data going somewhere it should not have — and it is categorically different from what this piece is about. The reason I open with it anyway is that it demonstrates the disclosure machinery firing on the lower-severity of the two categories, at a moment when the higher-severity one has produced no comparable public record. That absence is not reassurance. It is what the beginning of a category looks like.
Two populations, and only one of them is counted
The distinction that organises everything below is between a tool that is used and a system that acts. A shadow tool creates exposure when someone puts something into it: the failure mode is disclosure, the unit of exposure is the use, and the population is enormous. A shadow agent creates exposure when it does something: the failure mode is an action taken on the organisation's behalf, the unit of exposure is the action, and the population is small. Both are governance problems. They are not the same governance problem, and the standard practice of putting them on one slide called shadow AI is why most organisations are working the wrong one first.
The measurement asymmetry is stark and worth dwelling on. Shadow tool use is surveyed constantly. The PagerDuty survey conducted by Wakefield Research in 2026, across 1,250 office professionals at companies above half a billion dollars in revenue, found that two-thirds had used AI tools at work while believing they were not permitted to — a striking figure precisely because it measures knowing non-compliance rather than confusion. Nutanix's June 2026 financial-sector index, across 1,600 executives in fifteen countries, found that 86 percent believe AI used without official oversight creates business risk. There is no shortage of numbers.
For agents there is essentially one, and its shape is the finding. The same Nutanix survey reports that 66 percent of financial-services IT executives already encounter AI applications or agents introduced by employees outside IT. Read that phrase carefully: applications or agents. The best available published measurement of this problem cannot separate the two populations, which means nobody can tell an organisation how many acting systems it has. The number is not disputed. It does not exist.
Why tens rather than thousands is a claim about effort, and I want to be clear that it is a practice observation rather than a measurement. A person adopts a tool in ninety seconds: sign up, paste, go. A team builds an agent, which needs a model integration, a tool integration, credentials, somewhere to run and someone to keep it running. The barrier is real and it holds the population down. What it does not do is hold the exposure down, because exposure is the count multiplied by what each one can do — and the second factor is where an acting system differs from a chat window by orders of magnitude.
The difference is not severity, it is reversibility
It would be easy to read the argument so far as saying that agents are simply worse, and that is not quite the claim. The difference is structural, and it has three parts worth separating because each implies a different control.
The first is that an action is outward-facing. A leak is an internal fact about where data went, and however serious it is, the organisation can in principle bound it: enumerate what was exposed, notify, remediate. An action is a fact about the world. It has a counterparty who now believes something, and it may have created an obligation that a court will look at through the ordinary law of agency rather than through anything about AI. A system that emailed a customer a commitment, or approved a request, or amended a record another party then relied on, has done something the organisation cannot unilaterally take back — and the question of whether it had authority to do so is answered, in disputes with third parties, by what the counterparty reasonably understood rather than by what your internal approvals said. Apparent authority does not care that nobody signed off.
The second is that an agent repeats. A person who uses an unapproved tool badly does it a few dozen times before somebody notices. An acting system does the same thing several thousand times without variation, tiredness or hesitation, which means a scope error that a human would have caught by feel on the fourth iteration compounds silently. This is the same property that makes agents valuable, which is why the mitigation is scope rather than supervision.
The third is that an agent survives its author. A tool is used by a person; when that person leaves, the usage stops. An acting system keeps running, on a credential that frequently belongs to whoever built it, executing a set of assumptions that nobody currently employed can explain. Of everything in a shadow inventory, that is the category I would look at first: the systems whose owner column, once someone finally fills it in, names a person who left last year.
Four sources, and none of them is labelled agent
The reason a single scan never produces the inventory is that shadow agents arrive by four different routes, and each hides behind a different piece of organisational vocabulary.
Vendor-shipped, and usually the largest. The customer platform that acquired an AI assistant in a quarterly release. The service-management tool whose agent now closes tickets. The finance system that reconciles for you. These were procured by business units, approved through a process that assessed a product rather than a capability, and they are invisible to identity teams because their credential is the vendor's own integration. The discovery pattern is a pass over the vendor inventory looking for AI features shipped since the contract was signed, plus interviews with the business owners about what those features now do unsupervised. Most organisations find their largest acting population here, in software they believe they have already governed.
Team-built, and usually the second. The engineering team's triage agent, the finance team's reconciliation helper, the legal team's contract reviewer. Built with a team's own budget, its own cloud account, its own API key, and no central visibility because none was required. The discovery pattern is infrastructure-side: model API usage across cloud accounts, and credentials carrying write scope into systems of record. This is the category most people picture when they hear shadow agent, and it is not the biggest one.
Prototype-promoted, and the dangerous one. This category is created by nobody deciding anything. A hackathon project works, people start relying on it, reliance becomes dependence, and at some point it is load-bearing in production without a single moment at which anyone chose to put it there. It has the least governance, because it was never reviewed, and the longest operating history, because it has been running quietly for months. Discovery is behavioural rather than architectural: look in production for multi-step tool-call sequences that terminate in a write to a system of record. That signature is what an agent looks like from the outside regardless of what anyone calls it.
Integration-carried, the smallest and the hardest. The connector that was a data pipe and became an actor when the vendor shipped an update, described in release notes nobody reads, under a heading nobody flags. Nothing in the organisation's own systems changed. The discovery pattern is contractual — release notes and SaaS agreements reviewed for added AI capability — and it is unrewarding work that finds few things, some of which turn out to matter a great deal, which is a description of most useful security work.
The credential inventory is the harder half, and the more important one
Finding the agents is the part that feels like the exercise. It is the first half, and it is the easier half. The second half asks, for every acting system found, what credential it holds, where that credential lives, who else can reach it, when it was last rotated, and what it can actually do — and this half is harder because credentials are scattered across systems that no single team owns, and more important because the credential, not the agent, is what determines the blast radius.
The last of the five questions is the one that produces the uncomfortable findings, and it is the one most inventories skip because it takes real work. What a credential can do is not what the agent uses it for. Credentials are provisioned in a hurry, scoped generously because tight scoping breaks things during a build, and then never revisited, so the ordinary case is an agent doing three narrow things with a key that could do three hundred. Every shadow agent in your estate is, by definition, an accumulation of permissions with no entitlement behind them: something granted it the technical ability to act, and nobody with the authority to set its scope ever wrote down what it may do, because there was no moment at which anyone was asked.
The sizing context is worth giving properly, because it is routinely given badly. Published ratios of non-human to human identities in enterprises range from about 17:1 to about 144:1 depending on who is counting: Veza's 2026 identity and access work at the low end across a broader enterprise population; Rubrik Zero Labs at 45:1 for general enterprise; Entro Labs at 92:1 on a 2024 general-enterprise baseline and 144:1 for cloud-native and DevOps estates; Palo Alto Networks at 109:1 across nearly three thousand security leaders, of which 79 of the 109 are AI agents. Those numbers are frequently quoted against each other as though somebody must be wrong. Nobody is wrong. The denominators differ — a container-and-CI estate is not a typical enterprise — and the eight-fold spread is the finding, because it tells you the answer for your organisation is not knowable from anyone's report. It is knowable only by looking, and every one of those figures measures what one vendor's tooling could see, which makes each of them a lower bound.
Four dispositions, and the one everybody avoids
An inventory that produces no decisions is a document. For each acting system found, exactly one of four things has to happen, and the discipline is in forcing the choice rather than in the taxonomy.
- Keep. The system is valuable and is already governed well enough to stay as it is — a named owner, a scoped credential, a record of what it does. This is rare for anything discovered in a shadow inventory, and a keep decision reached quickly is usually a govern decision that nobody wanted to fund.
- Govern. The system is valuable, is not governed, and stays in production while the governance is built around it: a written entitlement, an enforcement point between its intent and the systems it writes to, and a record of what it did. This is the majority disposition and it is the one with a real cost attached, which is why it needs a named owner and a date rather than an intention.
- Demote. The system is valuable but is doing more than it should be trusted to do, so its scope is reduced until the governance catches up — from deciding to recommending, from write access to read, from production to pilot. Demotion is the most underused of the four and the most useful, because it preserves the value while removing the exposure, and it can be executed in a day.
- Decommission. The system is not worth the governance it would need. The credential is revoked, the infrastructure is removed, and — the part that is always forgotten — the credential's reach is traced and its access removed everywhere, not merely at the agent.
The avoided decision is demotion, and the reason is social rather than technical. Decommissioning is a clean verdict that a team can accept as a business call. Governing is a budget line that someone else pays. Demotion says, in front of the people who built it, that this thing works but is not currently trusted to act — and that is a harder sentence to say than either of the others. It is also the one that most often turns out to be right, because the value of most shadow agents is in the analysis they perform rather than the action they take, and the action was added because it was easy rather than because it was necessary.
Three markets, one inventory, three reasons it is asked for
United States: the inventory is the first thing requested and the last thing built. The supervisory position after 17 April 2026 makes this sharper rather than softer. The revised interagency model risk guidance — Fed SR 26-2 and OCC Bulletin 2026-13 — superseded SR 11-7 and placed generative and agentic AI expressly outside its scope, with separate guidance promised. Nothing about that deferral touches the duties attached to the action itself: the third-party risk expectations, the consumer-protection obligations and the safety-and-soundness duties all attach to what the system does, and a firm that cannot enumerate its acting systems cannot begin to demonstrate any of them. The CB Financial filing added a second, sharper reason: an examiner or a board that has read it now has a documented reason to ask what else is running, and the only acceptable answer to that question is a list.
India: the inventory is a data-processing question before it is a security one. Under the DPDP Act the relevant question about any acting system is not merely what it can do but what personal data it processes and for which stated purpose. A shadow agent is, in that framing, an unregistered processing activity — nobody declared its purpose because nobody declared its existence — and the exposure is not a disclosure event but a structural inability to answer a data principal's question or a Data Protection Board enquiry about a processing activity the organisation did not know it was performing. Indian organisations tend to find the credential half easier, because the identity estate is often younger and more centralised, and the purpose half far harder.
The Gulf: the inventory is named in the supervisory instrument. The Central Bank of the UAE's February 2026 guidance note expects licensed institutions to hold a risk-rated inventory of AI systems, and comparable SAMA expectations apply in the Kingdom. That is unusual and it is an advantage: the artifact is specified rather than inferred, which removes an argument that consumes months elsewhere. The corresponding risk is that an inventory built to satisfy a stated expectation gets built as a document rather than as a maintained register, and a register that is accurate on the day it is filed and stale a quarter later fails the only test that matters, which is being right when somebody asks.
Three weeks, and the control that has to land first
The inventory itself is bounded work, and I would run it in three passes rather than one. The first pass is identity-side: non-human identities that have taken an action in the last ninety days, filtered for model API access and for write scope into systems of record. The second is behavioural: production traffic scanned for the multi-step tool-call-then-write signature, which finds the systems whose credentials look ordinary. The third is conversational, and it finds what the first two structurally cannot — the newest systems, the most isolated ones, and the ones nobody has admitted to. Ask business units what they have built, what they have promoted, and what their vendors have shipped them, and ask it as an inventory question rather than a compliance one, because the same question asked the second way returns nothing.
The sequencing point matters more than the method. The control that prevents the next wave has to be announced before the inventory starts, not designed after it finishes, because an inventory announced without a control is an announcement of a deadline. Teams that hear a governance review is coming and have something they would rather not discuss will either accelerate to get in before it or go quieter, and both outcomes make the exercise worse. The control itself is unglamorous — no acting system reaches production without a review, the review has published criteria and a stated turnaround, and identity enforces it by not issuing a credential without the sign-off — and the only part that is hard is the first refusal, which is what makes the rest of it real.
There is a commercial argument here that is usually left implicit, and it deserves saying plainly. Organisations treat the inventory as overhead — the cost of being allowed to do the interesting work. It is the opposite. The inventory is the artifact a supervisor asks for first, the artifact a board committee needs before it can minute anything specific, the artifact due diligence asks for in a transaction, and the artifact that turns an incident from an open-ended question into a bounded one. The audit is the product. Everything else in an agent programme is a claim about capability; this is the only thing in it that is evidence.
What this argument does not prove
Four limits, and the first is the one I have been circling.
The tens claim is not measured. It is what I consistently see, it has a plausible mechanism behind it in the effort barrier to building an acting system, and it has no published census supporting it — because, as the diagram says, no such census exists. It is also actively eroding: agent-building frameworks have made the effort barrier much lower over the past eighteen months, and the honest forecast is that this becomes a hundreds problem within a couple of years, at which point the discovery methods described here stop scaling and something more like continuous detection is required. Anyone reading this in 2028 should check the premise before adopting the method.
Second, the CB Financial incident is doing rhetorical work in this piece that it cannot do evidentially. It is not an agent incident, I have said so, and the inference I draw from it — that the acting category will eventually produce comparable filings — is a forecast, not a finding. A reasonable person could read the same filing and conclude that the disclosure machinery works, that the failure was contained, and that the category is being handled. Third, the four sources and their ordering by size are a practice taxonomy rather than a measured distribution; the categories are useful because they imply different searches, but an organisation with a strong platform team will often find the second dominates, and the ranking should not be treated as data.
Fourth, and most substantively: an inventory is not a control. It tells you what exists on the day you looked. It does not stop anything, it does not govern anything, and it decays from the moment it is produced — which is why a one-off inventory delivered as a report is a worse outcome than a partial inventory delivered as a maintained register with an owner. The organisations that get value from this work are the ones that treat the first pass as the start of a register rather than the end of a project, and I have watched more programmes fail on that distinction than on any technical part of the exercise.
The question to take into the next risk committee is short and unusually diagnostic: how many systems in this organisation can take an action on our behalf without a human approving that specific action, and who owns each one? If nobody can answer it in a week, that is not a gap in the reporting. It is the finding, and it should be recorded as the finding rather than absorbed as a delay. If you are working out how to run that inventory before someone else asks you for it, that is the kind of engagement I do.