There is a slide I keep seeing in board packs, in some version, across every market I work in. It is near the back, it is titled something like AI update, and it carries a maturity chart with four bars and a green status marker. A capable executive presents it for eleven minutes. Somebody asks a good question about hallucination. The minute afterwards records that the board received an update on the company's use of AI and that no further action was required. Everyone in the room has discharged what they understood their obligation to be, and every one of them would be surprised to learn how little that page would be worth two years later.
Through the first half of 2026, a run of legal scholarship and practitioner commentary converged on a single claim: that AI governance has become a Caremark duty — that Delaware's law of board oversight now reaches the systems the company has built to make and take decisions. The convergence is genuine and it is worth taking seriously. It is also being repeated in a form that is imprecise in the two places where precision decides outcomes, and I want to fix both, because the imprecision is currently making boards complacent about the wrong things and anxious about the wrong things.
The first correction is deflationary: no Delaware court has adjudicated an oversight claim about artificial intelligence, the duty is not new, and the strongest of the 2026 scholarship argues explicitly that AI does not change the legal standard at all. The second correction runs the other way and is the reason this piece exists. The oversight duty is not a compliance standard. It sounds in loyalty, not care, which places it in exactly the category that the exculpation clause in your charter cannot touch — and that structural fact, rather than any new doctrine, is what makes an agent estate a personal problem for the people who signed off on it.
Caremark is not a compliance standard, and that is the part people skip
Start with the chain, because the load-bearing step is in the middle and it gets dropped in every summary I have read. In re Caremark, decided by the Delaware Court of Chancery in 1996, established that a sustained or systematic failure of board oversight can be a breach of fiduciary duty — while noting, in the same opinion, that it was possibly the most difficult theory in corporation law on which a plaintiff might hope to win. In Stone v. Ritter the Delaware Supreme Court in 2006 adopted the standard and characterised it: oversight liability is a failure to act in good faith, and the duty to act in good faith is a subsidiary element of the duty of loyalty rather than a free-standing duty of its own.
That characterisation is the whole game, for a reason that has nothing to do with AI. Section 102(b)(7) of the Delaware General Corporation Law permits a certificate of incorporation to eliminate directors' personal liability for monetary damages for breaches of the duty of care, and since the 2022 amendment permits the same for certain senior officers. Almost every Delaware corporation has such a clause, and most directors understand it, loosely, as protection against being sued for getting things wrong. The provision expressly does not permit exculpation for breaches of the duty of loyalty or for acts or omissions not in good faith. An oversight claim is a loyalty claim. The charter clause does not reach it.
The second load-bearing case is Marchand v. Barnhill, in which the Delaware Supreme Court in 2019 revived a dismissed oversight claim against the directors of a company whose central risk — food safety, in a business that made one product — had no board-level monitoring system at all. The holding that matters here is narrow and severe: where a risk is mission-critical, the board must have its own system for seeing it, and management's discretionary reporting up the chain is not a substitute for one. In 2023 the Court of Chancery extended a comparable oversight duty to officers within their own areas of responsibility. The doctrine has been moving steadily toward the position that the board must be able to see the thing itself, not a summary of the thing prepared by the people responsible for it.
Why AI plausibly clears the mission-critical bar, and why plausibly is the honest word
Mission-critical is not a defined term with a test attached; it is a characterisation a court applies after the fact, and the cases that have found it have generally involved a risk that was central to the business, heavily regulated, and obviously known to be so. Blue Bell made ice cream and the risk was listeria. Boeing made aeroplanes and the risk was that they fell out of the sky. Those are easy cases in hindsight, and the fact that they are easy in hindsight is exactly what should make a board uncomfortable, because nobody characterises their own risk as mission-critical in the minute before the incident.
For AI generally, I think the honest answer is that it depends on the company and that a lot of the commentary is overreaching. A retailer using a language model to draft product copy has not created a mission-critical risk and a court is unlikely to pretend otherwise. But the analysis changes sharply along one axis, and it is not the axis the commentary uses. It is not how much AI the company has. It is whether the AI acts.
A system that produces a recommendation which a person then adopts leaves the person in the chain of accountability, and the existing oversight architecture — which is built around people making decisions — continues to function. A system that takes the action itself removes the person and leaves the company committed by a process the board has never seen operate. Where that action is in a regulated function — a payment initiated, a claim decided, a limit adjusted, a customer communication sent, a record in a regulated system amended — the risk has the shape the cases care about: central, regulated, and known. My working position is that an agent estate with authority to act in regulated functions is the strongest available candidate for mission-critical characterisation in most regulated businesses, and that the same company's document-summarising copilots are not, and that treating those two as one AI risk is the analytical error most board reporting currently makes.
The 2026 convergence is commentary, not doctrine — which cuts both ways
Being precise about the status of the 2026 material is not pedantry; it changes what a board should do with it. The scholarship — most substantially Pierluigi Matera's work arguing that Caremark is not destabilised by artificial intelligence but activated by it — is academic argument. The Oxford Business Law Blog piece from March 2026 and the law-firm analyses that followed are practitioner guidance. The D&O commentary is an insurance market signal. None of it is a holding. As of August 2026, no Delaware court has ruled on an oversight claim arising from AI, and how Delaware fiduciary law will treat AI-specific oversight failures is genuinely unanswered.
Read carelessly, that is comfort. Read properly, it is the opposite, for two reasons. The first is that Matera's actual argument is more demanding than the headline version. He does not claim AI creates a new duty; he claims the doctrine remains one of loyalty rather than care, of bad faith rather than imperfect outcomes, and that what AI changes is how good faith has to be demonstrated. The board does not have to understand the internals of a model. It does have to show a sustained, documented effort to understand what the system does, to validate it, and to review it — and it has to do this without the crutch of understanding, which is the part boards find genuinely hard.
The second reason is that scholarship of this kind does not merely describe the law, it supplies the plaintiff's bar with its pleading theory. A derivative complaint has to survive a demand-futility motion on the documents, and the documents are the board's own. When the first AI oversight claim is filed — and the pipeline of public AI incidents makes it a question of timing — the complaint will be drafted against precisely this scholarship, and the defence will be the minutes. That is the practical relationship between commentary and doctrine, and it runs in the direction most boards are not preparing for.
There is a third element in Matera's analysis that I think is the most useful idea in the whole 2026 literature, and it has barely been picked up. Where the monitoring system is itself algorithmic, the red flags a board must respond to shift from the primary event to a secondary class: evidence that the monitoring system has stopped working — drift, degradation, false negatives. For an agent estate, generalise it one step further. The red flag is not the incident. The red flag is any indication that the organisation cannot tell whether an incident has occurred: an inventory that nobody can reconcile, a question about what an agent did last quarter that takes three weeks to answer, an internal audit finding that the logs do not support a reconstruction. Those are the board-visible symptoms, they arrive long before the incident does, and a board that receives one and does nothing has taken the specific step the doctrine punishes.
What the minute has to show
Everything above converges on a single artifact. The board's oversight is not what happened in the room; it is what the minute says happened in the room, because that is what a court will read years later with the outcome already known. This is the one place where the practice of governance and the law of governance are the same activity, and it is where most AI board reporting fails without anyone noticing.
The difference between the two columns is not diligence, it is specificity, and specificity is available only if the underlying material exists. A committee cannot minute that it reviewed the agent inventory as at a date unless someone maintains an agent inventory with a date on it. It cannot minute that it asked what the payment-initiating agents may do without human approval unless somebody can answer, which requires that entitlements exist as written objects rather than as accumulated permissions. It cannot minute that it noted two open findings unless internal audit has been given AI-scoped work. The minute is downstream of the plumbing, which is why boards that try to fix this at the minute-drafting stage produce documents that read, to anyone experienced, exactly like documents drafted to fix a problem at the minute-drafting stage.
One warning belongs here and it is not rhetorical. A minute recording oversight that did not occur is worse than no minute, because it is a document of the board's own making, in the plaintiff's hands, and the discovery that will accompany the claim will produce the underlying materials that do not support it. The instruction to a company secretary should never be to write better minutes. It should be to hold the meeting the better minute would describe.
The prong most boards have not implemented
If Marchand's requirement is a board-level system rather than management's reporting, then the practical question is what a board-level system for an agent estate could possibly consist of, given that everything the board knows about its agents arrives through the function being overseen.
Trace the chain and the problem is structural rather than cultural. The agent takes an action. The action leaves whatever trace the application was instrumented to leave. Engineering summarises for the business owner, who summarises for the executive, who reduces it to a slide with a status colour. At no point does anyone lie. At every point the information is compressed by someone with an entirely legitimate interest in the programme's continuation, and by the time it reaches the board it is a status rather than a fact. This is not a failure of integrity; it is what reporting chains do, which is precisely why the doctrine asks for something other than a reporting chain.
The something else is the fourth prong: independent verification. Not internal audit as it currently exists, which usually has no AI-scoped mandate, but a path by which someone who does not report into the programme can look at what the agents actually did and tell the committee whether it matches what the committee was told. That path has a hard prerequisite, and it is the reason this cannot be fixed in a board cycle: the underlying record has to have been written at the moment of action — what was done, on whose authority, against which policy, on what inputs — by a system built to write it. A board can direct a verification. It cannot conjure the evidence the verification would examine. If the records were never written, the independent reviewer is reduced to interviewing the same people the board was already hearing from, and the prong is decorative.
The supervisory deferral makes the board's own record the whole answer
There is a widespread and comfortable misreading of the April 2026 supervisory change that this argument has to confront directly. The revised interagency model risk management guidance issued on 17 April 2026 — Fed SR 26-2 and OCC Bulletin 2026-13, with the parallel FDIC issuance — superseded SR 11-7 and SR 21-8, and stated that generative and agentic AI models are novel and rapidly evolving and as such are not within the scope of the guidance, with separate AI guidance promised. It also declined to set enforceable standards or prescriptive requirements. A number of institutions have read that as breathing room.
It is not, and the reason is the one this practice keeps returning to. Out of scope is not out of risk. Every obligation attached to the underlying action — safety and soundness, consumer protection and fair lending, sectoral duties, third-party risk, and the ordinary corporate-law duty this piece is about — is untouched. What was withdrawn is the framework that would have specified the controls. A board that was hoping to discharge its oversight duty by pointing at a supervisory checklist has just been told there is no checklist, which means the board's own reasoning, its own inventory, its own questions and its own record are not a lesser form of evidence. They are the only form there is.
The asymmetry runs further than it first appears. When separate agentic guidance eventually arrives, it will be drafted against whatever the industry has by then built, in the way that supervisory expectations are always drafted against observed practice. The firms building serious oversight now are not merely protecting themselves; they are describing the baseline everyone else will be measured against. That is an unusual position for a governance function to be in, and it is the opposite of the compliance posture the topic usually attracts.
Three markets, three routes to the same obligation
United States: the doctrine is Delaware's, and the exposure is personal. For a US-incorporated public company the analysis is the one above, with one addition worth knowing about. The SEC's cybersecurity disclosure regime has already been triggered by an AI-rooted incident: in May 2026 CB Financial Services filed under Item 1.05 after an employee at its bank subsidiary used an unauthorised AI application to handle customer names, Social Security numbers and dates of birth. The materiality determination rested on the volume and sensitivity of the data, and the four-business-day clock ran from that determination rather than from discovery. That filing is a red flag in the doctrinal sense for every board in the sector — a publicly documented indication that the category produces material incidents — and a board that has since done nothing has a harder good-faith argument than a board that had never seen it.
India: the duty arrives through controls and committees rather than through case law. India has no Caremark, but it has a more prescriptive route to a similar place. Section 134(5)(e) of the Companies Act 2013 requires the directors of a listed company to state, in the Directors' Responsibility Statement, that internal financial controls have been laid down and were adequate and operating effectively — a signed assertion by named directors, made annually, about controls that an agent acting inside a financial process is now part of. SEBI's listing regulations require the top listed entities to constitute a Risk Management Committee with board membership, meeting at least twice a year, with a mandate that expressly covers cyber security. Layer the DPDP Act's obligations on the same estate and an Indian board's exposure is less about a derivative suit and more about a set of specific, dated, signed representations that an ungoverned agent estate quietly makes untrue.
The Gulf: board accountability is stated in the supervisory instrument itself. In the Gulf the obligation does not have to be derived from anything. The Central Bank of the UAE's February 2026 guidance note on AI adoption by licensed financial institutions expects a documented governance framework, a risk-rated inventory and board accountability for outcomes, and SAMA's frameworks in the Kingdom carry comparable expectations. Board accountability for outcomes stated in a supervisory instrument is, if anything, a sharper instrument than a fiduciary standard, because it is examinable rather than litigable and arrives on the supervisor's timetable rather than a plaintiff's. The structural advantage in the region is that much of the estate is new construction, so the inventory and the record can be designed in. The recurring weakness is that a documented framework and a governed estate are not the same artifact, and it is the second one an examiner eventually asks to see operating.
What to put in front of the board at the next meeting
Four things, in this order, and none of them requires a programme.
- The inventory, dated. Not a list of AI initiatives — a list of systems that can take an action on the company's behalf without a human approving that specific action, with an owner, the systems of record each can write to, and the classes of action each may take. If nobody can produce this, that is itself the finding, and it should be minuted as the finding rather than absorbed as a delay.
- The committee assignment, chartered. Where AI oversight sits — audit, risk, technology, or a dedicated committee — matters far less than that the charter says so, specifies what reporting the committee receives and at what cadence, and gives it the authority to commission its own advice. The charter is the artifact that demonstrates the board structured its oversight rather than encountered it.
- The two questions, asked on the record. What may each of these systems do without human approval, and what would we be able to produce if a regulator asked us in six months what one of them did last Tuesday? Minute the answers, including the ones that are we do not know, with a named owner and a date. An honest we do not know with a remediation date is a far better document than a confident assurance that discovery later contradicts.
- The independent look, commissioned. An AI-scoped internal audit or an external review that reports to the committee rather than through management, examining the same period the committee was briefed on and reporting whether the two accounts match. Once a year is enough. Never is the current default.
The cost of all four is a few days of executive time and one commissioned review, and the reason to do them at the next meeting rather than the next cycle is that the value of a minute is a function of its date. Oversight documented before the incident is evidence of a system. The same oversight documented after it is a response — and a court that is reading the minutes at all is reading them in the order they were written.
What this argument does not prove
Four limits, and I will start with the one that undercuts the piece hardest.
This may never be litigated. Caremark claims are extraordinarily difficult to plead and the great majority are dismissed; the doctrine requires bad faith, not negligence and not a bad outcome, and a board that held a mediocre discussion and minuted it thinly has a real argument that it made a good-faith effort. It is entirely possible that AI oversight failures are absorbed into ordinary securities and consumer litigation and that the derivative theory never becomes the vehicle. If that happens, most of the fiduciary framing in the 2026 commentary — mine included — will look overheated in retrospect. I still think the work is worth doing, because the artifacts it produces are the same ones a supervisor, an auditor and a plaintiff all ask for, but the strongest counter deserves to be stated plainly.
Second, I am an architect, not a lawyer, and nothing here is legal advice. The case law summarised above is settled and the exculpation point is a matter of statutory text, but the application of any of it to a specific board, a specific charter and a specific jurisdiction of incorporation is a question for counsel who has read the documents. Third, this is Delaware-centric by construction. A great many companies are not Delaware corporations, and the equivalent duties in other US states, in India, in the Gulf and elsewhere are differently shaped, differently enforced and differently insured. The three-market section above is a sketch of routes, not an equivalence claim.
Fourth, and most substantively: documented oversight is not good oversight. Everything in this piece is about making the board's engagement visible, and a board can produce an impeccable evidentiary record of a monitoring system that is monitoring the wrong things. The doctrine rewards the record because the record is what a court can assess, not because the record is what protects customers. I would rather a board asked one uncomfortable question badly minuted than four scripted ones perfectly minuted, and I have no way to reconcile that preference with the legal incentive except to name it.
The question that separates the two columns of the minute, and it is worth putting to your own board verbatim: if one of our systems acted wrongly last quarter, how would we find out, who would tell us, and how long would it take? If the answer runs through the team that built it, the board does not yet have a board-level system — it has a briefing. If you are working out what your committee should be receiving before someone else specifies it for you, that is the kind of session I do.