Pre-Mortem: The Liability Chain Medicare’s AI Prior Auth Model Has Not Drawn

On 1 January 2026, the Centers for Medicare and Medicaid Services in USA launched the WISeR model in six states, introducing prior authorisation to procedures that traditional Medicare had always provided without it. Contracted companies now assess medical necessity using AI. Human clinicians are required to sign off on any denial. The Senate voted 46-50 in July 2026 to keep the programme running. One question has not been answered.

This is the twelfth piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

CMS is wagering that AI-assisted prior authorisation reduces unnecessary Medicare spend without producing the patient-safety incident that forces a political reversal. If WISeR delivers measurable waste reduction without a documented causal chain from AI denial to patient harm, it becomes the template for prior authorisation across Medicare nationally. If it produces that chain, a documented line from AI recommendation to denial to patient harm, it does not just end WISeR. It becomes the reference point that makes AI prior auth politically untouchable in federal health programmes for a generation.

 

The Assumption

CMS has answered every operational question about WISeR except this one:

When an AI recommendation leads a contracted clinician to deny care and a patient is harmed as a result, where does liability sit?

The model design places a human clinician between the AI output and the denial decision. That establishes a paper trail. It does not establish a liability framework. Contractors earn between 10 and 20 per cent of the savings generated by denials and lose that payment when a denial is overturned on appeal. That is a commercial penalty, not a clinical one. The Federal Tort Claims Act does not cover contracted entities. No federal court has tested whether a contracted clinician reviewing AI recommendations at volume carries the same duty of care as a treating physician making an independent clinical judgement.

The assumption doing all the work in this model is that the human review layer is accountability enough. That assumption has not been tested.

 

The Sequence

1 July 2025. CMS published the WISeR notice in the Federal Register and did not submit it to Congress under the Congressional Review Act. That omission would matter later.

1 January 2026. WISeR launched in New Jersey, Ohio, Oklahoma, Texas, Arizona, and Washington.

17 March 2026. The Washington Post published an exclusive: Medicare’s new AI gatekeeper was delaying care for seniors. The University of Washington’s medical system had nearly 100 patients waiting for epidural injections. In Arizona, Phoenix pain specialist Dr Matthew Crooks told Medscape that every epidural injection submitted in the first three months had been denied and described the system as completely nonfunctional and unsustainable. In Texas, initial AI approval rates ran at 62 per cent, against a 92 per cent national approval rate across Medicare Advantage.

25 March 2026. The Electronic Frontier Foundation filed a FOIA lawsuit against CMS in federal court in California, seeking records on WISeR’s AI algorithms, training data, bias safeguards, and the financial incentives paid to contractors. The suit confirmed that CMS had not made its AI methodology or vendor compensation structure publicly available seven weeks after launch.

6 April 2026. CMS published a Federal Register notice delaying prior authorisation implementation for certain services within the model to allow additional time for operational readiness. CMS also issued a corrective action order against one of its AI contractors. Both confirmed that the model’s operational design had not performed as intended in the first quarter.

12 May 2026. The Government Accountability Office issued its determination: WISeR met the Administrative Procedure Act definition of a rule and was subject to the Congressional Review Act. CMS had not made the required submission to Congress before the model took effect.

20 May 2026. Senator Ron Wyden and Representatives Suzan DelBene and Greg Landsman introduced resolutions of disapproval in both chambers, seeking to repeal WISeR under the CRA.

6 July 2026. Gold carding launched in Washington state. Providers achieving a 90 per cent affirmation rate across a minimum of ten prior authorisation requests become exempt from further review for covered services. Quarterly rollout to the remaining five states is planned.

16 July 2026. The Senate voted 46-50 against advancing the disapproval resolution. Party line. WISeR survived. The liability question the GAO had exposed survived with it.

The Pager

Dr Mehmet Oz, Administrator of the Centers for Medicare and Medicaid Services.

The message: WISeR’s accountability chain has not been drawn. The model places a contracted clinician between an AI denial recommendation and a Medicare beneficiary, but no published document establishes where negligence sits when a patient is harmed following an AI-assisted denial. The Federal Tort Claims Act does not cover contractors. Contractors point to the human clinician. Clinicians are reviewing AI output under volume pressure with no published duty-of-care standard for that specific context. When the first federal lawsuit tests this configuration, and one will, CMS will need a published framework, not a contract clause. That framework is easier to write before litigation than after.

 

The Proof

Gold carding is the model’s self-correction mechanism. If quarterly rollout reaches all six states and the 90 per cent affirmation threshold functions as a genuine quality signal, the AI layer contracts over time as trust is established. Proven providers exit prior auth. New entrants face the review. The model becomes calibrated rather than blanket.

If gold carding stalls or rollout criteria are applied inconsistently across jurisdictions, the AI layer expands without a release valve. Prior auth burden accumulates regardless of provider track record. The model becomes a cost-reduction instrument with no exit for providers who have earned one.

The proof of the bet is not the aggregate savings figure. It is whether WISeR, by the end of 2026, has published a liability framework and delivered gold carding in all six states. Without both, the model is running on the same untested assumption it started with.

 

Verdict

If CMS publishes a liability framework for AI-assisted denials before a federal case forces the question, and gold carding delivers consistent rollout across all six states, WISeR will be the strongest government evidence yet that AI-assisted utilisation review can reduce Medicare waste without a patient-safety crisis. The accountability design would become the reference for every federal health programme that follows.

Without the liability framework, WISeR accumulates its risk quietly. Not through a single dramatic incident, but through the gap between AI recommendation volume and human review capacity, compounded by an accountability vacuum no published document has yet closed. That gap does not stay open indefinitely.

Pre-Mortem: The US Government’s 3,611 AI Use Cases

On 3 April 2025, the White House issued OMB Memorandum M-25-21, directing every major federal agency to appoint a Chief AI Officer, expand the use of artificial intelligence across government operations, and manage risk proportional to each system’s impact on citizens. Twelve months later, the Federal Agency AI Use Case Inventory records 3,611 AI use cases across 56 agencies, more than double the prior year’s total. A May 2026 survey of more than 200 technology executives across civilian and defence agencies found 53% are actively planning agentic AI pilots. Only 8% of those agencies have incident response frameworks in place.

This is the eleventh piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

The US government is betting that embedding AI across 3,611 federal workflows covering benefits decisions, immigration adjudications, healthcare determinations, and law enforcement, will make government faster and more efficient before the accountability architecture governing those decisions is clarified. OMB M-25-21 requires Chief AI Officers, public AI inventories, and risk management proportional to impact. The hard compliance deadline for role-specific AI training arrives in September 2026. If that architecture catches up to the deployment before a consequential wrong decision reaches a citizen with no named relief, the bet holds.

 

The Assumption

The expansion’s credibility turns on one unanswered question: whether the Federal Tort Claims Act, designed to govern negligent acts by human federal employees, applies without amendment to decisions made by AI agents running inside federal systems. The same May 2026 survey found only 44% of agencies include vendor liability clauses in AI contracts, and only 29% have documented kill-switch procedures. The legal architecture governing accountability in federal government was designed for humans acting on behalf of the state. No court has ruled on whether it extends to the agents they built.

 

The Sequence

In 2024, federal agencies reported 1,757 AI use cases. By 2025, that figure had grown to 3,611. In March 2026, the Department of Veterans Affairs expanded AI use in claims processing, with 215 of its 367 AI systems classified as high-impact, covering benefit eligibility, healthcare access, and fraud detection. In May 2026, the majority of agencies were planning agentic pilots, with only 20% having defined pre-deployment testing policies. The AI reached citizens before the accountability reached the AI.

 

The Pager

Russell Vought, Director of the Office of Management and Budget, carries the M-25-21 mandate at the centre of the federal AI expansion. Every covered agency has designated a Chief AI Officer responsible for inventory, risk management, and AI governance at agency level. The VA alone runs 215 high-impact AI systems. No single published document names what relief is available to a veteran whose claim was influenced by one of those systems, which official carries accountability for that decision, or whether the Federal Tort Claims Act applies when the acting party is software, not a civil servant.

 

The Proof

The measure that would settle this is a published legal standard: a named accountability chain clarifying who carries liability when a federal AI agent makes a consequential wrong decision, whether government, vendor, or joint, and whether the Federal Tort Claims Act applies or new legislation is required. No such standard has been published. OMB M-25-21 requires risk management proportional to impact. It does not name the relief available to a citizen when that risk management fails, nor the date by which that question must be answered.

 

Verdict

If OMB publishes, before the September 2026 training compliance deadline, a named accountability standard for AI-driven decisions in high-impact federal systems, covering who carries liability when the AI is wrong and what legal remedy a citizen holds, the expansion will stand as the most deliberate attempt the US federal government has made to govern AI before it reaches citizens at scale. Without that, the US government has put AI into 3,611 workflows and left the question of who carries the call when the AI gets it wrong to be answered in court, by accident, or not at all.

Pre-Mortem: The UK’s Critical Third Party Regime

On 13 July 2026, Amazon Web Services, Google Cloud, Microsoft Azure, and Oracle became the first companies formally designated as Critical Third Parties to the UK financial system. The Bank of England, the Prudential Regulation Authority, and the Financial Conduct Authority now hold powers to gather information, assess resilience, and make enforceable rules against the four providers for the services they supply to the financial sector. A 2024 Bank of England and FCA survey found the top three cloud providers accounted for 73% of all cloud providers named by respondents across the UK financial sector. The designation names the risk. It does not resolve it.

This is the tenth piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

The UK is betting that direct regulatory oversight of four technology providers, applied specifically to their financial-sector services, will reduce the systemic risk from having most of the sector’s cloud infrastructure concentrated in three companies. The Financial Services and Markets Act 2023, which created the CTP regime, gives the Bank of England, PRA, and FCA powers to assess resilience and enforce CTP-specific rules. The designation is a supervisory relationship, not a structural remedy. If that supervisory relationship produces documented, published improvements in resilience before the first major cloud incident in UK financial services, the bet holds.

 

The Assumption

The regime’s credibility turns on one scoping decision: that overseeing four providers for the services they supply to the UK financial sector is sufficient to contain risks generated by four companies whose infrastructure decisions are made globally, across legal jurisdictions and customer bases far larger than the UK financial system. Microsoft Ireland Operations Limited is the designated entity. Its architecture decisions are made in Redmond. The supervisory perimeter covers the financial-sector slice. The concentration risk does not stop there.

 

The Sequence

The concentration risk pre-dated the regime by years. The Financial Services and Markets Act 2023 established the legislative basis for the CTP framework. A 2024 Bank of England and FCA survey confirmed the scale: three providers controlling the majority of UK financial-sector cloud infrastructure. HM Treasury announced the first four designations on 10 July 2026, effective 13 July. The sequence is legislation, then evidence, then designation. The risk was present throughout.

 

The Pager

Rachel Blake MP, Economic Secretary to the Treasury and City Minister, made the designation announcement. The Bank of England, PRA, and FCA share oversight of the four providers under the regime. Three regulators. Three separate mandates. No published document names which of the three leads incident coordination when a designated provider’s outage affects UK financial services. The CTP framework assigns supervisory responsibility. It does not assign the call.

 

The Proof

The measure that would settle this regime’s effectiveness is a published resilience outcome: a before-and-after comparison of systemic vulnerability at a named date after the CTP rules take effect. No such commitment has been published. The three regulators hold powers to gather information from the four providers. No public document names what information will be published, in what form, and by when. The first formal review cycle has no published date.

 

Verdict

If the three regulators jointly publish a named lead for CTP incident coordination and commit to a quantified resilience outcome before the first formal review cycle, the designation will stand as the most substantive step the UK has taken to address cloud concentration risk in its financial sector. Without that, four of the world’s most powerful technology companies have been formally named, and the framework that names them has not yet named who is in charge when one of them goes down.

Pre-Mortem: The EU AI Act’s Accountability Gap


On 2 August 2026, the EU AI Act gives the EU AI Office the power to fine the developers of general-purpose AI models up to three per cent of global annual turnover, demand documentation, and commission independent access to source code. Three weeks before that date, the high-risk AI compliance deadline moved from August 2026 to December 2027, enacted as binding law on 29 June. The two facts share a date. They do not share a plan.

This is the ninth piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

The EU is betting that extending the deadline for high-risk AI compliance by 16 months, agreed in May 2026 and enacted on 29 June, produces better enforcement outcomes than a met deadline inside a half-prepared enforcement architecture. The logic holds. As of August 2026, only nine of 27 member states have shown advanced public implementation of enforcement infrastructure. Germany has designated the Bundesnetzagentur as its market surveillance authority and adopted draft transposition legislation. Spain built the AESIA, a dedicated supervisory agency, from scratch. Ireland deployed fifteen coordinated authorities under a central National AI Office. Those are genuine structural commitments. Eighteen member states have not reached that point. The extension gives them time. Whether they use it is the bet.

 

The Assumption

The entire framework rests on this: that national competent authorities, operating under 27 different legal frameworks, will converge on consistent enforcement before December 2027. The AI Act is a directly applicable regulation. Its enforcement infrastructure is not. The regulation sets the rules uniformly across the bloc. The authorities responsible for applying them have been built at very different speeds, under very different political conditions. That divergence is the risk the extension is buying time to close. There is no public commitment that the time is sufficient.

 

The Sequence

The AI Act entered into force in August 2024. Member states were required to designate their national competent authorities by August 2025. At least twelve missed that deadline. Seven months later, in May 2026, the Council and Parliament agreed to simplify the rules as part of the Digital Omnibus package. On 29 June, the high-risk AI deadline moved. What remains in force on 2 August is a narrower set: general-purpose AI model obligations and transparency requirements for new deployments. The high-risk AI rules, the Act’s original centre of gravity, are no longer in that set. Governance was adjusted to fit the readiness gap. That is not the order in which enforcement architecture is supposed to be built.

 

The Pager

Lucilla Sioli, Director of the EU AI Office, carries accountability for general-purpose AI enforcement from 2 August. For high-risk AI systems, including credit-scoring models, recruitment tools, and systems used in border control, healthcare, and law enforcement, accountability rests with national competent authorities. In 17 of 27 member states, no public designation exists. The Act names the category. Seventeen member states have yet to name the person.

 

The Proof

The measure that would settle this in 2028 is year-one enforcement consistency: the share of member states that have conducted at least one formal high-risk AI enforcement action, under the same evidentiary standard, in the first twelve months after the December 2027 deadline. No EU institution has publicly committed to publishing that figure. The AI Office’s annual progress reporting is the closest mechanism on the public record. It tracks activity. No published mechanism commits to measuring whether enforcement actions are consistent across member states.

 

Verdict

If the Commission designates a public accountability owner in each member state before December 2026 and commits to publishing year-one enforcement data by name, the 16-month extension holds up as a governance decision made under realistic conditions. Without that, a framework that took two years to reach enforcement hands itself an extension with nobody carrying it.

Pre-Mortem: Eight Companies, No Published Accountability Standard

The Pre-Mortem is a weekly series on this blog. Each piece applies five questions to a major technology commitment before the outcome is known.

In February 2026, the United States Department of War signed agreements with eight of the world’s leading artificial intelligence companies, OpenAI, Google, Microsoft, SpaceX, Oracle, Amazon Web Services, NVIDIA, and Reflection, to deploy their advanced AI models inside its classified networks. Impact Level 6 (IL6) covers data classified at the Secret level. Impact Level 7 (IL7) covers compartmented intelligence and the most sensitive operational systems, where the United States military runs its actual warfighting decision support. This is the first time that large language models have operated within IL7 environments. What has not been published is who carries accountability when one of them gets something wrong.

 

The Bet

The Department of War’s stated aim is to establish the United States military as an AI-first fighting force, achieving what its AI Acceleration Strategy calls decision superiority across all domains of warfare. The eight agreements are the mechanism. The AI systems will summarise surveillance feeds, synthesise intelligence data, and suggest tactical options to human operators. The Department of War’s five AI ethics principles, responsible, equitable, traceable, reliable, and governable, are on the record. The bet is that those principles are sufficient architecture for what happens inside a classified environment.

 

The Assumption

The whole bet turns on this: that “humans remain accountable for AI outcomes” as a stated principle is equivalent to a published accountability framework.

That distinction is where there is a gap. The Department of War’s Responsible AI Strategy and Implementation Pathway establishes process. It does not name the specific individual, command role, or governance layer accountable when an AI-assisted intelligence summary inside an IL7 environment shapes a decision that turns out to be wrong. Principle and framework are not the same thing, and in a classified environment that distinction cannot be tested publicly.

 

The Sequence

In July 2025, Anthropic’s Claude became the first frontier AI model approved for use on classified networks. The Pentagon subsequently sought to renegotiate those terms, demanding Anthropic permit its models to be used for all lawful purposes without limitation. Anthropic declined, citing concerns about mass domestic surveillance and autonomous weapons. On 27 February 2026, President Trump ordered all federal agencies to stop using Anthropic. The following day, OpenAI signed its classified deal with commitments that included prohibitions on domestic mass surveillance and human responsibility for the use of force, positions that aligned with the guardrails Anthropic had sought to retain. By May 2026, the remaining seven of the eight, Google, Microsoft, SpaceX, Oracle, Amazon Web Services, NVIDIA, and Reflection, had signed equivalent agreements.

The sequence reveals something structural. The accountability architecture for classified military AI was settled by commercial negotiation and political designation, not by a published governance framework.

 

The Pager

Legal scholars on autonomous weapons identify the same accountability fracture that applies in the decision-support context here. When an AI-assisted output causes harm in a classified environment, accountability distributes: software developers could not have anticipated all operational contexts, commanding officers disclaim responsibility for machine-generated outputs, vendors invoke contractual limitation of liability. The human-in-the-loop design means a person reviews AI suggestions before acting. It does not mean accountability for acting on a wrong AI output has been named anywhere in the command chain.

No published document names the specific individual role, command layer, or governance body accountable for a wrong AI-assisted output inside an IL7 environment. No congressional oversight mechanism covers classified operational AI use. No published error reporting standard exists. By the nature of classified operations, none can.

 

The Proof

Eight companies, the highest classification levels, large language models operating on top-secret data for the first time: the scale of the commitment is confirmed. The outcome data will not follow. Classified operational AI performance is not publicly reviewed, by design. This is the only deployment in this series where the proof question cannot be answered from the outside, not because the data is not collected, but because it cannot be published.

The accountability question is not whether humans are in the loop. They are, by stated commitment. The question is whether the framework for who carries it specifically, when they get something wrong, inside a system that cannot publish what it got wrong, exists in any enforceable form.

 

The Verdict

If the Department of War’s five principles are operationalised into a named, enforceable command accountability chain for AI-assisted decisions at every classification level, if the commercial guardrails in all eight agreements are independently verifiable by a body with appropriate clearance, and if a congressional oversight mechanism specific to classified AI operational failure is established, then this is what responsible military AI deployment at scale should look like.

Without all three, eight of the most powerful AI systems on earth are running inside the most classified networks in the world. The decisions they shape will not be publicly reviewed. The wrong ones will not be counted.

The accountability is a principle. The framework has not been built yet.

Pre-Mortem: Apple Intelligence at Work

The Pre-Mortem is a weekly series on this blog. Each piece applies five questions to a major technology commitment before the outcome is known.

On 9 June 2026, Apple used its annual developer conference to announce that Siri had become something different. Not a smarter assistant. An agentic AI layer that could take actions across applications, services, and workplace workflows on behalf of its users, across a hardware ecosystem of more than 2.5 billion active devices. The world’s most valuable company had turned its operating system into an AI agent. The question the keynote did not answer was straightforward: when it gets something wrong at work, who is responsible?


The Bet

Apple is betting that privacy and accountability are the same problem. Its Private Cloud Compute architecture is genuinely novel: stateless, ephemeral, cryptographically auditable, with production builds published within 90 days for independent inspection. At WWDC 2026, Craig Federighi stated: “data is only used to execute your request, and outside experts can continue to verify this promise at any time.” The claim is that if Apple cannot read your data, no one can. What this architecture was not designed to answer is what happens when Apple Intelligence takes a workplace action on your behalf and gets it wrong. That is a different question. Apple has framed the privacy answer as if it covers both.


The Assumption

Everything turns on one distinction: that an architecture designed to prove Apple cannot access your data also constitutes a framework for enterprise accountability when AI actions produce incorrect outcomes.

It does not. Privacy means Apple is not the party reading your data. Accountability means someone is responsible for what the AI produces from it. Those are different obligations. No document currently published by Apple closes the gap between them. The existing AppleCare for Enterprise terms explicitly disclaim liability for lost profits, damage, corruption, or loss of data, or interruption of business. There is no AI-specific carve-out, no enterprise service level agreement for Apple Intelligence outputs, and no accuracy standard committed to publicly.


The Sequence

Three weeks before WWDC 2026, Apple settled a $250 million class action over Siri AI features it had promoted during the iPhone 16 launch but did not deliver. The settlement included no admission of wrongdoing. In April 2026, Apple’s CEO Tim Cook announced his departure from the role, with John Ternus, the head of hardware engineering, confirmed as his successor from September 1, 2026. Ternus had no publicly stated role in shaping Apple Intelligence. At WWDC 2026, enterprise MDM controls for Apple Intelligence were available in beta only, with general availability expected in autumn 2026. The agentic deployment was announced. The governance controls that enterprises need to deploy it responsibly were not yet generally available.


The Pager

Craig Federighi, Senior Vice President of Software Engineering, is the named face of Apple Intelligence. Amar Subramanya, Vice President of AI, is the operational lead, reporting to Federighi since the retirement of John Giannandrea earlier this year. Neither has made any public commitment regarding enterprise accountability for AI outputs. By September 2026, John Ternus will carry the CEO accountability for a deployment he did not architect, operating under governance terms that were written before agentic AI was part of the product. No named individual or governance body is publicly committed to what Apple Intelligence does in enterprise workflows when it goes wrong.

The Proof

Apple has published no enterprise outcome measure for Apple Intelligence. No accuracy benchmark, no error rate commitment, no service level agreement for business customers. The company’s transparency commitments for Private Cloud Compute are real: production code published within 90 days, a cryptographically auditable log, a virtual research environment for security testing. These are privacy verification mechanisms, not performance standards. A survey of approximately 100 enterprise IT administrators published in May 2026 found that the primary concern was data exfiltration to unmanaged providers, and that eight per cent of organisations had already moved to prohibit AI features entirely. No one at Apple has publicly committed to a measure that would settle that question.

The Verdict

Apple has done more than most technology companies to make its cloud AI architecture independently verifiable. Private Cloud Compute is a credible attempt to resolve the privacy half of the enterprise AI problem. The accountability half remains open. If Apple publishes enterprise terms that define who carries responsibility for agentic errors in business workflows, and if John Ternus names a specific accountable owner for enterprise AI governance before the full iOS 27 rollout, the MDM controls announced at WWDC 2026 become the foundation of something credible. Without both, the hundreds of millions of Apple Intelligence-enabled devices deployed into enterprise settings are operating on a privacy promise. That is not the same thing as an accountability framework.

Pre-Mortem: KPMG’s AI-Powered Audit

The audit opinion is the most consequential document most public companies produce. Not the annual report. Not the investor deck. The audit opinion, because it carries a named partner’s signature, and because that signature means something in law. On 9 June 2026, KPMG and Microsoft announced the deployment of Microsoft Agent 365 and Copilot across 276,000 KPMG professionals in 138 countries, including inside KPMG Clara, the firm’s global smart audit platform. Scott Flynn, KPMG’s Global Head of Audit, called it “a pivotal milestone in our AI-powered, human assured audit transformation.” The word “assured” is doing a great deal of work in that sentence.

A pre-mortem asks the same five questions, every time, applied before failure is possible rather than after. This is the fifth in the series. The first looked at vendor accountability in regulated finance. The second at clinical safety in healthcare. The third at execution accountability in defence procurement. The fourth at clinical AI infrastructure. This one looks at professional services, the sector that has built its entire business model on the premise that human expertise is the product.

 

The Bet

KPMG is betting that efficiency and accountability can coexist at this scale. That 276,000 professionals deploying AI agents, with a governance layer running underneath, will not dilute the professional accountability the audit opinion rests on. It is a reasonable bet. It is also an untested one. The commercial logic is clear: 276,000 professionals, 138 countries, and an AI-powered workflow running through KPMG Clara creates the kind of structural productivity gain that redefines the firm’s cost base, and potentially its fee model. Analysis of recent audit fee movements suggests clients are already pressing the case that AI efficiency should flow through to lower fees. The deeper bet, the one sitting beneath the headline deployment, is that “AI-powered, human-assured” constitutes a defensible operating model before any regulatory body has defined what “human-assured” actually requires in practice.

 

The Assumption

The single assumption carrying all the weight: that governing agents is the same thing as being accountable for them. Microsoft Agent 365 provides what its own documentation describes as a control plane, a centralised registry of agents with lifecycle rules, identity controls, and audit logging. That is a meaningful capability. It answers the question: how many agents do you have, and what can they touch? It does not, on its own, answer the question a claims lawyer or a regulator will eventually ask: who is accountable when the agent was visible, governed, and still wrong? KPMG’s Trusted AI framework lists ten ethical pillars, including one labelled Accountability, which calls for human oversight and responsibility to be embedded across the AI lifecycle. That is a principle-level commitment. None of the publicly available documentation specifies what happens to the partner’s signature when an AI-assisted conclusion is signed off and later found to be materially incorrect.

 

The Sequence

KPMG has deployed agents at scale before any authoritative regulatory framework specifies what AI-assisted audit evidence must look like, or how human review of AI-generated conclusions must be documented to meet existing standards. The IAASB approved a project proposal in March 2026 to revise ISA 500, Audit Evidence, to address technology use in audit, but the project is still in early research and information gathering, with no exposure draft issued and no effective date. The PCAOB has stated publicly that it is considering developing risk management guidance for audit firms using AI. Considering, not publishing. The capability is deployed. The standard that surrounds it is still being drafted.

 

The Pager

Lisa Heneghan, KPMG’s Global Chief Digital Officer, was specific about what this deployment requires: “strong foundations in governance, visibility and accountability.” That framing is responsible, and Agent 365 provides the visibility that most enterprises currently lack. The harder question is structural and specific. The audit opinion is signed by a named partner. Professional indemnity is priced around that signature. When an agent embedded in KPMG Clara surfaces a conclusion, the partner reviews it, signs the opinion, and the work later contains a material error, the liability has historically sat with the partner and the firm. What KPMG, Microsoft, and the client have not yet published is a clear allocation of responsibility for the agent’s contribution to that error. Is it a tool failure, an oversight failure, or something existing frameworks do not yet classify? The governance layer provides the audit trail. It does not specify who reads it, or what reading it is worth, when a claim is filed.

 

The Proof

The announcement commits 276,000 professionals and earns KPMG the designation of Microsoft “Frontier Firm.” Neither is a performance measure. No published metric connects this deployment to audit accuracy improvement, reduction in deficiencies, or quality outcomes. What the deployment actually demonstrates is that KPMG can deploy Agent 365 at scale and maintain visibility over its agent estate. That is a meaningful operational achievement. It is not the same as demonstrating that AI-assisted audit conclusions are more reliable than human-only ones, which is what regulators, courts, and insurers will eventually need to see. KPMG Clara’s existing framing covers adoption and workflow integration. No published figure connects it to audit opinion accuracy or deficiency rates. The proof that matters most is still outstanding.

 

Verdict

If KPMG publishes a clear framework specifying how AI-assisted audit evidence is reviewed, validated, and documented, paired with a liability position that survives regulatory scrutiny, this becomes the reference model for professional services AI at scale. The governance commitment is genuine. The scale of deployment is unmatched in the sector. Scott Flynn’s “AI-powered, human-assured” is the right aspiration. The question is whether “human-assured” describes a documented, auditable review process that a regulator will accept and an insurer will cover, or whether it is a positioning statement waiting for a definition. At 276,000 professionals across 138 countries, the audit opinion at the centre of this deployment is too consequential to leave that question open. The answer should come before the first material claim, not after.

Already Building: Epic Agent Factory and the Governance Gap

The pre-mortem on Epic Agent Factory asked who would answer when a health-system-built agent made a clinically significant error. It published on 9 June. I have since learned of a Becker’s Hospital Review report from 30 March confirming that one of America’s largest health systems had already been building those agents for weeks before the question was published.

It confirms the pre-mortem’s central argument. Neither the research nor the article surfaced how quickly the sequence had already begun.

 

The Deployment That Was Already In Motion

Advocate Health had already tapped Epic’s Agent Factory, becoming one of the first health systems to build and deploy agents through the platform. Andy Crowder, Advocate Health’s SVP and Chief Digital and AI Officer, described the direction in a LinkedIn post on 26 March: “By combining Epic’s Agent Factory Platform capabilities with Advocate Health’s scale, clinical insight, and commitment to innovation, we’re translating AI from promise into practice.” He pointed to a three-day Epic immersion at The Pearl innovation district in Charlotte, focused on speeding up pharmacy verification for complex medications and cutting infusion chart preparation time for pharmacists and nurses. Four working prototypes emerged, scheduled to go live in July 2026.

Crowder added: “Together, we’re advancing responsible, practical AI that fits naturally into clinical workflows, reduces friction, and gives clinicians back time to focus on what matters most.” It is a considered statement, and the commitment is genuine. But it is not a governance document. And Advocate Health is not unusual here. They are representative. They moved first because the platform enabled it, the commercial pressure to reduce administrative burden was real, and nothing in the regulatory landscape said stop.

This is the sequence the pre-mortem described. Capability arrived. Deployment followed. The governance architecture to surround it had not been ratified.

 

The Workflows That Come Next

Pharmacy verification and infusion chart preparation are not, in themselves, clinical decision-making. They reduce documentation burden and carry genuine operational value. But they are the entry point, not the ceiling.

Epic’s own Penny agent already handles prior authorisation for thousands of health systems. Agent Factory is the platform through which health systems build their own versions of exactly those capabilities. Prior authorisation sits at the intersection of clinical judgment and payer approval. An AI-generated argument that misrepresents a contraindication, omits a relevant diagnosis, or positions a clinical case in a way that leads a payer to deny appropriate care causes harm that is downstream and deniable. The agent did not make the clinical decision. But the agent shaped the argument that influenced it.

The pre-mortem’s central question, who owns the error, was always pointed at this trajectory. The agent is built by the health system, on Epic’s platform, using Curiosity’s foundation models, in a regulatory environment where no one has yet specified how liability is allocated between vendor and deployer. Advocate Health’s prototypes are the first step of a sequence that leads directly to that question.

 

Colorado Tried to Build the Rails

While health systems were building, legislators in Colorado were attempting to create the governance scaffolding that the platform lacks at a federal level. Three separate AI-related healthcare laws had been passed by June 2026, each addressing a different dimension of the problem, and each confirming the same underlying gap.

Colorado’s original AI Act, SB 24-205, was scrapped before it ever took effect. A legal challenge from X.AI in April 2026, supported by federal intervention from the DOJ, led to enforcement being suspended and the legislature repealing the law entirely. Its replacement, SB 26-189, was signed on 14 May. It is a narrower law, retaining consumer notice requirements and the right to meaningful human review following adverse outcomes, but dropping the duty-of-care standard and mandatory impact assessments that had made the original controversial. It takes effect January 1, 2027.

HB 26-1139, signed on 2 June, constrains how payers use AI in coverage determinations. It requires that AI-driven decisions be based on the patient’s individual medical and clinical history rather than group data, and that any denial or delay of coverage based on medical necessity receive review by a licensed clinician. It too takes effect January 1, 2027.

Together, SB 26-189 and HB 26-1139 create obligations on both sides of the prior authorisation workflow. Neither specifies who bears the cost when an agent-generated output leads to the wrong clinical outcome. Three laws confirming the gap exists is not the same as closing it.

 

The Sequence Is Not a Prediction. It Is a Pattern.

On 1 June 2026, eight days before the pre-mortem was published, the Joint Commission launched its first voluntary AI certification programme for healthcare organisations. Built on the initial guidance published with the Coalition for Health AI in September 2025, the certification covers governance, data management, risk and bias reduction, and monitoring. It is a meaningful step forward. But the certification recognises organisations, not individual tools. It does not validate or certify individual AI products. It contains no discussion of liability allocation. It is a framework for responsible intent, not a mechanism for accountability when something goes wrong.

Epic has not published a liability framework specifying what a health system owns when a self-built Agent Factory agent produces a clinical error. No Epic contract language or public terms of service document does so. No federal regulatory body has published guidance specifically addressing liability allocation for agentic AI operating within EHR environments. The FDA has authorised more than 1,400 AI-enabled devices and issued no specific enforcement guidance for agentic AI in EHR environments.

The pre-mortem’s conclusion was that if Epic published a clear liability framework and paired it with a safety review mechanism, Agent Factory could become the defining infrastructure layer of hospital AI over the next decade. That conclusion stands. What the evidence now confirms is that the clock is not running from some future launch date.

It was already running.

Pre-Mortem: Epic Agent Factory

Update, 14 June 2026: One of America’s largest health systems was already building Agent Factory agents in late March, weeks before this piece published. This new piece confirms the central argument.


 

Epic unveiled Agent Factory at HIMSS 2026 (March 2026), positioning it as a no-code, drag-and-drop visual builder that lets health systems design, deploy, and monitor their own autonomous AI agents inside the Epic environment. Alongside it came Curiosity, a family of generative medical foundation models trained on deidentified records from 300 million patients across 310 health systems, backed by a research preprint on arXiv first published in August 2025. Together, the announcements represent Epic’s move from AI vendor to AI infrastructure provider, handing health systems the tools to build clinical automation at their own pace and on their own terms.

A pre-mortem is a discipline borrowed from project risk management. Before a programme succeeds or fails, you ask: if this does not go as planned, what was the mechanism? This series applies that lens to major AI-in-industry announcements, not to predict failure but to surface the questions that deserve answers before deployment, not after.

 

The Bet

Epic is betting that health systems want to own their AI destiny. Phil Lindemann, VP of Data and Research, framed Agent Factory as enabling customers to implement AI solutions without needing to call a vendor or write a line of code. That is a significant commercial and philosophical shift. Epic’s existing suite, Art, Penny, and Emmie, has posted credible numbers: 42 per cent reduction in prior authorisation submission time at Summit Health, 58 per cent sustained reduction in billing-related service messages at Rush University, 69 per cent early lung cancer detection at The Christ Hospital against a 46 per cent national average. The bet is that health systems, given those results as proof of concept, will want to build the next generation themselves.

 

The Assumption

The assumption underneath Agent Factory is that health system capability is ready to meet platform capability. Canvas Medical CEO Adam Farren noted in HIMSS 2026 commentary that most hospitals are not yet positioned to take advantage of the platform. Agent Factory is in early phase, with first availability in 2026 and continued rollout in 2027. Epic’s own roadmap, and the organisational readiness required for clinical agent deployment, put realistic momentum at leading health systems two to three years out. The platform may well be sound. The question is whether the organisations it serves have the clinical informatics depth, the governance infrastructure, and the project bandwidth to build and validate autonomous agents safely, particularly in clinical rather than administrative workflows.

 

The Sequence

Epic shipped the capability before any ratified standard governs what happens when a health-system-built agent makes a clinically significant error. The Joint Commission and Coalition for Health AI published voluntary joint guidance in September 2025, covering governance structures and vendor management. The FDA has authorised over 1,400 AI-enabled devices but has published no specific enforcement guidance for agentic AI in EHR environments. No federal regulatory framework yet specifies how liability for agent-generated clinical errors should be allocated between vendor and deploying health system. The capability is real and available. The governance architecture to surround it is not yet ratified.

 

The Pager

When an Agent Factory-built agent makes a clinically significant error, who owns it? Epic’s public framing places health systems “in the driver’s seat.” That is a positioning statement, not a governance document. No published contract language, terms of service excerpt, or named executive statement specifies who bears liability for agent-generated errors. No Epic accountability framework for self-built agents has been published. KPMG’s Q4 AI Pulse Survey (2025) found that 75 per cent of large-enterprise leaders name security, compliance, and auditability as their top requirements for agent deployment. At present, the answer to the pager question is that nobody has publicly claimed the call.

 

The Proof

Curiosity carries published research behind it: a preprint on arXiv first submitted in August 2025, covering 118 million patients and 151 billion tokens via the CoMET architecture. That is a meaningful evidential bar. Agent Factory has no equivalent published validation. Epic’s self-reported statistic that more than 85 per cent of customers are actively using Epic AI is plausible given market penetration of 43.7 per cent of US hospitals by count and 56.9 per cent by beds, but it refers to the existing suite, not to Agent Factory specifically. No performance benchmarks, error rate thresholds, or clinical outcome commitments for health-system-built agents on Agent Factory appear in any public source.

 

Verdict

If Epic publishes a clear liability framework that specifies what health systems own when they deploy self-built agents, and pairs that with a safety review mechanism before clinical agents go live, Agent Factory could become the defining infrastructure layer of hospital AI over the next decade. The foundation is genuinely strong: real outcome data from deployed agents, a clinically substantiated foundation model, and a market position that no competitor can easily replicate. The Curiosity publication demonstrates that Epic is capable of meeting an external evidential standard. The question is whether it applies that same rigour to the governance scaffolding around Agent Factory before health systems start building in earnest, rather than after the first serious incident forces the issue.

Pre-Mortem: The Pentagon’s Autonomous Drones Reset

 

The Pentagon’s Replicator programme promised thousands of cheap autonomous drones in two years and delivered hundreds. The response has not been to wind it down. It has been to dissolve it, rebuild it as a new command inside Special Operations Command, and ask Congress for roughly 240 times the money. A programme that under-delivered on a lean, fast model is being re-attempted on a vast one, and the case for why the second structure succeeds where the first did not has not yet been made in public.

A pre-mortem asks the same five questions, every time, applied to a current programme before failure is possible rather than after. This is the third in the series. The first looked at vendor accountability in regulated finance. The second looked at clinical safety accountability in regulated healthcare. This one looks at execution accountability in defence procurement, the hardest delivery environment of them all. Different sector, similar structural shape: commitment moving faster than the architecture meant to hold it to account.

 

The Bet

The bet is that scale fixes what speed could not. Replicator was announced in August 2023 with a target of multiple thousands of all-domain attritable autonomous systems inside roughly two years, run by the Defense Innovation Unit on about a billion dollars across two fiscal years. It was deliberately lean, built to route around the traditional acquisition machine. By the deadline it had fielded hundreds. The reset, the Defense Autonomous Warfare Group, carries a 2027 budget request of about $54 billion, against roughly $226 million the year before. The technical bet is sound on its face: mass autonomy is where warfare is going, and the United States cannot afford to be slow to it. The harder bet, the one sitting under the headline number, is that money and a command structure fix what was an execution problem. Those are different things, and the launch treats them as one.

 

The Assumption

One belief is doing all the work: that Replicator’s shortfall was a problem of resourcing and structure, solvable with more of both. The documented failures point elsewhere. Systems were selected that proved unreliable, too expensive, or too slow to manufacture at the quantities needed. Some existed only as a concept when they were chosen. And the programme could not procure software able to orchestrate and command large, mixed swarms of different drones, which is the actual technical heart of autonomy at scale. None of those is a budget problem. A bigger budget buys more of the same systems and more of the same integration gap. If the diagnosis is wrong, the cure scales the disease.

 

The Sequence

Commitment came before the architecture, again. Replicator launched in August 2023. A second line of effort, focused on countering small drones, was added by a Secretary of Defense memo in September 2024. The original thousands-by-2025 deadline arrived with hundreds delivered. The programme was then consolidated into a joint interagency task force, dissolved, and rebuilt as the new autonomous-warfare group inside Special Operations Command, with the first acquisition under the new structure landing in January 2026, two counter-drone systems. Only in April 2026 did the Secretary tell the House Armed Services Committee that a sub-unified command for autonomous warfare was coming. The command meant to own this is still being stood up around a commitment already made. The funding tells the same story. Of that $54 billion, only about $1 billion is appropriated base money. The other $53 billion is a request, parked in a flexible five-year reconciliation pot that Congress has not yet passed. The headline number signals overwhelming commitment. In hard terms it is roughly a billion dollars in hand and fifty-three billion in hope. The intention is real. The money, for now, is one dollar in every fifty-four.

 

The Pager

Start with the credit, because it is real. The new group has a named director, Lt. Gen. Francis L. Donovan (USMC), with a clear command line and an appointment made by the Secretary himself. That is more named, senior accountability than most large defence programmes ever put on the public record, and it counts for something. The harder question is operational and specific. Standing policy requires appropriate levels of human judgement over the use of force. At swarm scale, with attritable systems acting at machine speed, who is the named individual accountable when one of them engages wrongly? The command line is clear. The accountability for the autonomous decision itself, at the scale this programme is built to reach, has not been framed in public. A command answers for a programme. It is a harder thing to say who answers for a single autonomous engagement when there are thousands of them in the air.

 

The Proof

The committed measures are input measures. Dollars requested, units contracted, the first systems bought. There is no public outcome measure for capability actually delivered, no cost per effective intercept, no fielded-and-working-at-scale figure with a date attached. This matters because the proof problem already bit once. Leadership called Replicator on track in 2024 and said it had made enormous strides in 2025, while the independent accounting found hundreds, not thousands. When the people who own the programme also own the definition of progress, optimism outruns delivery. Second-attempt scepticism is earned, not unfair. In eighteen months, the question of whether this worked will be answered by whoever holds the platform to define what delivered at scale means, and right now that platform is a budget request.

 

Verdict

This is a serious programme with serious people behind it. The strategic logic is correct, mass autonomy matters and slowness is its own risk. The accountability has a name and a rank, which is rare. The first systems have been bought and are heading to the field. None of that is in doubt.

What is unproven is whether a command and a budget can fix a problem that was about manufacturing maturity, software orchestration, and realistic system selection. A reorganisation addresses none of those by itself.

The action is concrete. Publish the outcome measure, not the input: a fielded-and-working-at-scale metric with a date, committed before the reconciliation money is spent, not after. Name the human accountable for autonomous engagement decisions at scale, not only the command that owns the programme. And diagnose the first shortfall in public before scaling, so the much larger second bet rests on a corrected understanding rather than a hope.

If the department publishes a delivered-at-scale outcome measure tied to a named owner, and solves the swarm-orchestration software problem it could not solve the first time, this becomes the programme that proves autonomous capability can be fielded at speed. Without both, it becomes the most expensive way yet found to relearn that money and reorganisation do not fix an execution problem.