The Fiduciary Blind Spot: Caremark Requirements When AI Systems Change


Caremark does not require directors to prevent every AI failure. But it may require an information-and-reporting system capable of detecting consequential changes after deployment and escalating them to someone who can act.  

The hardest oversight question presented by adaptive AI is not whether directors can understand a model’s source code. It is whether the corporation has a reliable way to learn when a consequential system has materially changed, determine whether the change matters, and place the issue before someone with authority to act.

A recent Delaware case involving overdraft fees, not AI, offers a useful analogy. In Brewer v. Turner, the court allowed a shareholder derivative suit to proceed against more than a dozen current and former directors of Regions Financial Corporation, the parent of Regions Bank. The bank had working groups, an audit committee, and outside counsel. But a whistleblower complaint describing allegedly unlawful overdraft-fee practices reached the board in November 2019, and the board’s response, hiring a law firm to investigate, did not lead to corrective action for roughly a year and a half. During that period, the bank continued to collect the fees. Regions ultimately paid $191 million to settle claims brought by federal regulators. Chancellor Kathaleen McCormick concluded that the delay in responding to a known red flag, rather than the absence of a reporting system, was sufficient to allow the derivative claim to proceed.

That case previews the fact pattern that adaptive AI may create for boards. The issue may not be that a company lacked an oversight structure. It may be that a signal, such as a drift alert, an anomalous output pattern, or an internal warning, reached the organization, but no one with authority to act responded quickly enough.

That possibility fits within Delaware’s existing oversight doctrine. Caremark does not require directors to prevent every operational failure or forecast every model error. It requires a good-faith effort to establish a reasonable information and reporting system, and a good-faith response when that system produces warning signs. For software that can be retrained, updated, connected to external tools, or deployed with limited human intervention, a reasonable reporting system must account for what happens after launch, not only what was approved before it.

This does not mean that every inaccurate output creates fiduciary liability. Nor does it mean that every board must establish an AI committee or review model weights. The narrower point is practical: when adaptive software is central to a company’s safety, compliance, or operations, a governance program that stops at deployment may fail to monitor the risk the company actually faces.

Caremark Standard in an Adaptive Environment

The modern discussion begins with In re Caremark International Inc. Derivative Litigation and Stone v. Ritter. Together, those decisions describe two related oversight theories: a board’s failure in good faith to attempt to establish a reasonable reporting system, and a board’s conscious failure to monitor an existing system or respond to known red flags. The standard is demanding. Oversight claims generally require particularized facts supporting an inference of bad-faith disloyalty, not merely a showing that a control could have been better or that a business decision turned out badly.

Marchand v. Barnhill illustrates why the company’s business matters. In that case, the Delaware Supreme Court allowed a claim to proceed where a complaint alleged that the board of a single-product ice-cream manufacturer had no board-level process for receiving food-safety information, even though food safety was central to the company’s business. The decision did not require directors to guarantee that contamination would never occur. It focused on whether the board had made a meaningful effort to ensure that a mission-critical risk reached the boardroom.

 

Later decisions, including In re Clovis Oncology and In re Boeing, likewise show the importance of information flows concerning risks at the center of a company’s operations. Those decisions should not be read as eliminating the high bar for Caremark claims. They do, however, make clear that a general compliance program will not necessarily answer a claim involving a risk that is both mission-critical and poorly reported.

 

The same logic applies to adaptive systems. A scheduling tool used for ordinary administrative work does not present the same oversight problem as software that changes network access, allocates hospital resources, balances an electric grid, or generates recommendations that employees or customers may rely on. The relevant question is not whether a product is marketed as “AI.” It is whether the system’s behavior can materially affect the company’s legal or operational exposure and whether that behavior can change without a conventional approval event.

 

The distinction also matters for officers. The Court of Chancery’s decision in In re McDonald’s Corp. Stockholder Derivative Litigation recognized that officers may owe oversight duties within their areas of responsibility. That principle connects the technical work of a Chief Information Security Officer, Chief Technology Officer, Compliance Officer, or business-unit leader to the board’s reporting system. A board may delegate execution. It cannot assume that delegation eliminates the need for escalation, ownership, and follow-up.

 

Why Pre-Deployment Compliance Is Not Enough

Traditional software often produces familiar operational signals: outages, error codes, failed transactions, or unauthorized access. Adaptive systems can deteriorate more quietly. Concept drift occurs when the relationship between inputs and outcomes changes. Data-distribution shift occurs when production conditions diverge from those represented in testing. Feedback loops can make both problems harder to detect because the system’s decisions alter the environment that supplies its future data.

 

A clean launch, successful red-team exercise, or vendor attestation establishes a baseline. It does not establish that the baseline will remain valid six months later. A model may be retrained on new data, connected to a new application programming interface, exposed to a new population, or placed in a workflow where users begin relying on it in ways the original design did not anticipate. The resulting risk may arise from the interaction of several individually approved changes rather than from a single defective line of code.

 

For oversight purposes, the key issue is translation. Management should be able to explain what the system does, where it acts, which inputs it depends on, what it is permitted to change, what signals indicate deterioration, and who must be notified. The board need not evaluate the model’s internal weights. It should be able to ask whether the company has defined the system’s purpose and boundaries, tested its principal failure modes, retained a record of material changes, and assigned authority to pause or roll back the system.

 

Vendor dependence makes this problem more acute. A company may bear responsibility for decisions made with a third-party tool while lacking access to the tool’s training data, update history, or performance telemetry. A contract that addresses only uptime and service levels may not address model drift, unexplained output changes, audit rights, incident notification, or the vendor’s obligation to support a rollback. Those omissions do not automatically establish a fiduciary breach. They do create a foreseeable information problem if the system is important enough that the board would reasonably expect management to monitor its legal and operational effects.

 

The practical lesson is not that companies must demand disclosure of every proprietary detail. Procurement and oversight should be aligned. If a vendor cannot provide model-level transparency, the company may need compensating controls: outcome monitoring, independent validation, restrictions on autonomous action, periodic reapproval, or a contractual right to suspend the service when defined thresholds are crossed.

 

Operationalizing Oversight Across Infrastructure, Health, and Power Sectors

The stakes are clearest where automated decisions have immediate legal, physical, or financial consequences.

 

In cybersecurity, an automated defense system may isolate devices, revoke credentials, or act against infrastructure outside the company's network. A false positive can disrupt a partner, and an unauthorized countermeasure can create civil or criminal exposure. The board's job is not to approve each rule; it is to confirm management has set authorization boundaries, escalation thresholds, and a tested kill switch.

 

In health care, an intake or triage model may behave differently during a surge or demographic shift, with risks spanning patient safety, discrimination, privacy, and medical-device obligations. The question is whether the organization can catch a materially different pattern of recommendations before it becomes routine.

 

In electricity and freight, adaptive tools that reallocate loads or route shipments can cascade even when each recommendation looks plausible on its own. The board need not design the fallback; it should confirm one exists, has been tested, and has a named owner.

Across sectors, the failure is usually cumulative, not catastrophic: small shifts in data or user behavior move software away from validated conditions before anyone notices. The gap is rarely a failure to predict an error; it is a failure to assign someone to notice the drift.

A Useful Regulatory Analogy: Predetermined Change Control Plans

The FDA framework for Predetermined Change Control Plans (“PCCPs”) offers a useful, though limited, analogy for thinking about adaptive AI governance. The analogy is not that FDA device regulation establishes corporate-law requirements for every company using AI. Rather, it is that both frameworks must address how an organization can authorize and monitor changes to a system whose behavior may evolve after deployment.

 

The FDA’s 2025 final guidance describes a PCCP for an AI-enabled device as a plan that identifies the modifications the manufacturer intends to make, the methodology for developing, validating, and implementing them, and an assessment of their effect. The agency can review the plan as part of a marketing submission, allowing specified changes without requiring a new submission for each modification. A PCCP is not a corporate-law safe harbor, and the FDA’s device framework cannot simply be imported into enterprise governance. Its relevance here is structural: it establishes predefined boundaries and evidence for authorized change.

 

That structure rejects two unhelpful extremes. Freezing all updates is commercially unrealistic, while permitting unconstrained adaptation is difficult when a system has safety or compliance consequences. The corporate-governance lesson is therefore narrower than “companies must adopt PCCPs.” It is that companies should determine in advance which changes are permitted, what evidence must support them, and which changes require review or escalation.

 

A company can apply that logic internally. Before deployment, management can define a permitted change envelope: the kinds of retraining allowed, the data sources that may be added, the metrics that must remain within range, the actions that remain subject to human approval, and the conditions that require rollback or escalation. The plan should also state what evidence will demonstrate that a change remains within the envelope. That evidence might include validation results, outcome comparisons, drift measures, override rates, incident reports, and time-to-revert.

 

The purpose is not to turn an FDA device-control concept into a universal corporate-law mandate. It is to create a governance record connecting the system’s intended use to the company’s monitoring and intervention capabilities. When the system changes, the company should know whether the change was authorized and tested, whether it remained within the approved boundaries, and who reviewed the result.

 

Five Questions for Legally Meaningful Oversight

A practical board-level program can be organized around five questions.

 

First, which systems have mission-critical consequences? Management should maintain an inventory that classifies systems by potential impact on safety, legal compliance, operations, customers, employees, and material assets. Product labels are less useful than the consequences. A model described as a “recommendation engine” may still be mission-critical if personnel routinely follow its output without independent review.

 

Second, what level of autonomy and reversibility does each system have? Risk depends not only on accuracy. It also depends on the system’s authority, speed, access to external tools, ability to affect third parties, and the difficulty of reversing its actions. A model that makes a low-stakes suggestion presents a different oversight problem from an agent that can change access rights or execute transactions.

 

Third, which changes may occur without escalation? The company should specify what can change automatically and what requires review. Data-source additions, retraining outside validated conditions, material shifts in error rates, or changes to the system’s authority should not be left to informal judgment.

 

Fourth, what evidence will reveal deterioration? Useful telemetry should be tied to decisions and consequences, not limited to technical uptime. Relevant indicators may include error and override rates, unexplained output shifts, performance across affected groups, drift measures, incident frequency, and time-to-revert. Independent review is particularly important where the business unit benefiting from automation also evaluates whether the automation is working.

 

Fifth, who can intervene, and how quickly? A reporting system is incomplete if no person has both technical ability and legal authority to pause, investigate, roll back, or deactivate a high-consequence system. The answer should be documented, tested, and available outside ordinary business hours when the system operates continuously.

 

Allocating Responsibility and Avoiding Overclaiming

The intersection between internal board oversight and public disclosure obligations under SEC rules adds a second layer of accountability. Since 2023, Item 106 of Regulation S-K has required public companies to describe, in their annual reports, their processes for assessing and managing material cybersecurity risk and the board's role in overseeing it. The rule does not by itself extend to AI governance, but it establishes the template regulators are likely to apply as adaptive-system risk becomes a more familiar subject of Form 10-K disclosure: a description of process, not merely an assurance of confidence.

SEC v. SolarWinds shows both the promise and the limits of disclosure litigation. In 2024, the Southern District of New York dismissed most of the SEC's claims against SolarWinds and its Chief Information Officer, including the theory that deficient cybersecurity practices violated the securities laws' internal-accounting-controls provisions. The court held that the provision applies to financial accounting, not cybersecurity generally, and rejected claims based on hindsight about post-breach disclosures. But it allowed a claim based on a specific pre-breach security statement to proceed because the company allegedly maintained it despite contrary internal evidence. The lesson is narrow: generalized governance assurances are unlikely to create liability, but a specific representation about oversight can do so when the company knows it is inaccurate.

Conclusion

The Caremark question for adaptive AI is not whether directors can guarantee that a model will remain accurate. It is whether they have made a good-faith effort to ensure that material changes in a consequential system will be detected, understood, and reported to someone who can act.

 

That standard leaves room for experimentation and recognizes the limits of board-level technical expertise. It also rules out a governance model in which responsibility ends at launch. A dashboard without an owner is not oversight. A rollback plan that has never been tested is not a control. And a vendor certification that says nothing about post-deployment change cannot substitute for an information system capable of detecting when the company’s risk has changed.

 

For boards, the immediate task is modest but concrete: identify the systems that matter, define their permitted boundaries, monitor evidence of deterioration, and name the person authorized to intervene. The fiduciary blind spot is not that AI changes. It is failing to build a reporting system that notices when it does.