Pre-Mortem: The UK’s Critical Third Party Regime

On 13 July 2026, Amazon Web Services, Google Cloud, Microsoft Azure, and Oracle became the first companies formally designated as Critical Third Parties to the UK financial system. The Bank of England, the Prudential Regulation Authority, and the Financial Conduct Authority now hold powers to gather information, assess resilience, and make enforceable rules against the four providers for the services they supply to the financial sector. A 2024 Bank of England and FCA survey found the top three cloud providers accounted for 73% of all cloud providers named by respondents across the UK financial sector. The designation names the risk. It does not resolve it.

This is the tenth piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

The UK is betting that direct regulatory oversight of four technology providers, applied specifically to their financial-sector services, will reduce the systemic risk from having most of the sector’s cloud infrastructure concentrated in three companies. The Financial Services and Markets Act 2023, which created the CTP regime, gives the Bank of England, PRA, and FCA powers to assess resilience and enforce CTP-specific rules. The designation is a supervisory relationship, not a structural remedy. If that supervisory relationship produces documented, published improvements in resilience before the first major cloud incident in UK financial services, the bet holds.

 

The Assumption

The regime’s credibility turns on one scoping decision: that overseeing four providers for the services they supply to the UK financial sector is sufficient to contain risks generated by four companies whose infrastructure decisions are made globally, across legal jurisdictions and customer bases far larger than the UK financial system. Microsoft Ireland Operations Limited is the designated entity. Its architecture decisions are made in Redmond. The supervisory perimeter covers the financial-sector slice. The concentration risk does not stop there.

 

The Sequence

The concentration risk pre-dated the regime by years. The Financial Services and Markets Act 2023 established the legislative basis for the CTP framework. A 2024 Bank of England and FCA survey confirmed the scale: three providers controlling the majority of UK financial-sector cloud infrastructure. HM Treasury announced the first four designations on 10 July 2026, effective 13 July. The sequence is legislation, then evidence, then designation. The risk was present throughout.

 

The Pager

Rachel Blake MP, Economic Secretary to the Treasury and City Minister, made the designation announcement. The Bank of England, PRA, and FCA share oversight of the four providers under the regime. Three regulators. Three separate mandates. No published document names which of the three leads incident coordination when a designated provider’s outage affects UK financial services. The CTP framework assigns supervisory responsibility. It does not assign the call.

 

The Proof

The measure that would settle this regime’s effectiveness is a published resilience outcome: a before-and-after comparison of systemic vulnerability at a named date after the CTP rules take effect. No such commitment has been published. The three regulators hold powers to gather information from the four providers. No public document names what information will be published, in what form, and by when. The first formal review cycle has no published date.

 

Verdict

If the three regulators jointly publish a named lead for CTP incident coordination and commit to a quantified resilience outcome before the first formal review cycle, the designation will stand as the most substantive step the UK has taken to address cloud concentration risk in its financial sector. Without that, four of the world’s most powerful technology companies have been formally named, and the framework that names them has not yet named who is in charge when one of them goes down.

Healthcare AI Enters Its Accountability Phase

Healthcare AI has stopped being an experiment. Holland & Knight, the US law firm, put it plainly in its mid-2026 healthcare report: the sector has entered a “recalibration phase,” where capital discipline and demonstrable return on investment have replaced the growth-first logic that funded the last five years of digital health.

That is a legal and investment framing, not a clinical one. But the clinical evidence backing it up is now specific enough to name.

Kaiser Permanente’s Permanente Medical Group rolled out ambient AI scribing to 7,260 physicians across more than 2.5 million patient encounters between October 2023 and December 2024. The result, confirmed by Kaiser’s own Division of Research: nearly 16,000 clinician-hours of documentation time saved. Not a pilot cohort. Not a vendor’s projection. A production deployment, measured after the fact, across a workforce large enough that the number means something.

Ambient documentation is also the part of healthcare AI with the least room left to argue about. A 2026 survey of 120 US health systems, run by the healthcare research firm Eliciting Insights, found clinical note-taking and ambient listening tools now sit at 68% adoption, up 62% year on year. Among the health systems able to quantify results, 61% report at least a 2x return specifically from ambient listening tools.

 

The $3.20 Figure Is Real, and Older Than It Looks

The oft-quoted “$3.20 return for every $1 invested in healthcare AI” is genuine, but it is worth knowing where it actually comes from before repeating it in a board pack. It traces to a Microsoft-sponsored IDC study published in late 2023 and reported in early 2024, not a fresh 2026 finding. It has simply become the industry’s standing benchmark figure, cited so often across 2025 and 2026 coverage that it now reads as current data. It is not wrong. It is just two years old and vendor-commissioned, which matters if you are the one deciding how much weight to put on it.

The Kaiser and adoption figures matter more, precisely because they are recent, specific, and independently reported rather than recycled.

 

What Actually Produced the Return

This ROI happened because of a specific programme design, not because someone bought a good tool, one that most other sectors experimenting with AI have not adopted.

Kaiser did not deploy ambient scribing and then discover the workflow around it. Clinical documentation workflow got redesigned first, and the AI tool was the mechanism, not the starting point. Accountability for the outcome, hours saved, adoption sustained, clinician trust maintained, was established before rollout, not retrofitted afterwards to justify the spend. And the whole exercise operated under exactly the capital discipline Holland & Knight describes: prove the return, or the funding does not continue.

That sequence, workflow redesign first, accountability from day one, capital discipline over growth optimism, is the actual explanation for why healthcare produced verifiable ROI while most other sectors are still producing pilot decks.

 

Healthcare Is Now the Benchmark, Not the Exception

Treat healthcare’s result as evidence that AI works and you will draw the wrong lesson. The technology was never really in question. What was in question, and what most other sectors are still failing to answer, is whether the organisation deploying it redesigned anything before switching it on.

Healthcare had no choice but to answer that question properly. Clinical documentation errors have consequences that show up in patient outcomes and malpractice exposure, not just quarterly numbers, so the sector could not afford the deploy-first governance-later approach that has quietly become normal everywhere else.

That is what other sectors should actually be benchmarking against: not whether their AI produces a return, but whether their programme was ever designed to make one provable.

 

The Question Worth Asking Before the Next AI Business Case

Before signing off the next AI investment, the question is not whether AI delivers value. Healthcare has already answered that question, under specific and now well-documented conditions.

The real question is whether your programme has been designed to match those conditions, workflow redesign before deployment, accountability defined from the outset, capital discipline over growth optimism, or whether it has been designed the way most digital health investment was designed before 2026: fund it, hope the outcomes show up eventually, and find out later whether anyone was ever going to check.

Healthcare already found out. That is the whole difference.

The Programme Succeeded. That Was the Problem.

Most digital transformation programmes are designed to end.

That is the actual design flaw. Not the technology chosen, not the budget allocated, not even the ambition behind the initiative. The programme itself is structured around a finish line, a go-live date, a steering committee sign-off, a moment when the work is declared complete and the team disbands.

The problem is that the market, the technology, and the customer never agreed to stop moving at that point.

 

The Project Mindset Is the Actual Liability

Transformation programmes are built like construction projects. Define the scope, execute the plan, hand over the keys, move on to the next thing. That structure works well for building a bridge. It works badly for building an organisation’s capacity to keep adapting, because the moment the programme ends, so does the organisation’s active attention to the problem it was meant to solve.

A May 2026 Forbes Business Council analysis puts it plainly: treating digital transformation as a project sets the expectation that there is a finish line to cross. There is not. Markets keep moving. Customer expectations shift faster than any single programme can track. Data environments and operating models change shape well after the sign-off. A transformation programme with a defined end date is optimised for a world that stopped changing the day the programme closed, which is not the world any organisation actually operates in.

The Forbes analysis draws a comparison that holds up well: digital transformation works like fitness. When you stop, you atrophy. Nobody who has kept fit for a decade did it with a single twelve-week programme and then stopped. They built a habit that never formally ends.

 

What Continuous Capability Actually Looks Like

Tesla is the clearest large-scale example of what this looks like in practice. Tesla ships software improvements to vehicles already on the road through over-the-air updates, rather than treating the car’s capability as fixed at the point of sale. Autopilot and Full Self-Driving features are refined through frequent releases, often tested on a small subset of vehicles before wider rollout, rather than waiting for a full model cycle to bundle every improvement together. The car’s capability keeps changing for as long as the vehicle is on the road, never declared finished at any single point.

Most organisations do not need to ship software to a fleet of vehicles. But the underlying pattern is the same: small changes shipped continuously and tested before wide release, rather than large changes bundled into an infrequent big-bang release. That pattern is exactly what separates organisations still adapting years after their transformation programme closed from the ones still running the same processes the programme was meant to replace.

 

The Five Things That Actually Change

Shifting from a transformation mindset to a continuous one is not about abandoning structure. It requires deciding to do a small number of things differently, and doing them consistently.

Start with the conversation itself. The goal is not to convince stakeholders that transformation was wrong. It is to convince them that the current approach stops too early. Frame the shift as doing transformation properly, not as replacing it with something else.

Build modular, not monolithic. Large, all-or-nothing platform overhauls are exactly the kind of investment that locks an organisation into a single technology decision for a decade. Modular, scalable components can be replaced or upgraded individually as needs change, without requiring another multi-year programme to do it.

Treat learning as infrastructure, not an event. A single training push before go-live does not build a capability. Continuous training, embedded guidance, and space to experiment safely are what actually let people keep pace with a system that keeps changing.

Change what gets measured. Tracking project completion tells you the programme finished. It tells you nothing about whether the organisation can still adapt six months later. Track agility, the rate of continuous improvement, and customer outcomes instead, because those are the metrics that actually describe ongoing capability.

Build the feedback loop permanently. Regular input from employees and customers is not a phase of the programme. It is the mechanism that tells the organisation when the next adjustment is needed, and it only works if it never switches off.

 

The Question Worth Asking Before the Next Transformation Sign-Off

Before the next transformation programme gets a steering committee sign-off and a closing date, the honest question is whether the organisation’s capacity to keep adapting exists independently of the programme that is about to close, not whether the scope was delivered.

If the answer is no, the programme did not fail to transform the organisation. It succeeded at exactly what it was designed to do, and the design was the problem.

Pre-Mortem: The EU AI Act’s Accountability Gap


On 2 August 2026, the EU AI Act gives the EU AI Office the power to fine the developers of general-purpose AI models up to three per cent of global annual turnover, demand documentation, and commission independent access to source code. Three weeks before that date, the high-risk AI compliance deadline moved from August 2026 to December 2027, enacted as binding law on 29 June. The two facts share a date. They do not share a plan.

This is the ninth piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

The EU is betting that extending the deadline for high-risk AI compliance by 16 months, agreed in May 2026 and enacted on 29 June, produces better enforcement outcomes than a met deadline inside a half-prepared enforcement architecture. The logic holds. As of August 2026, only nine of 27 member states have shown advanced public implementation of enforcement infrastructure. Germany has designated the Bundesnetzagentur as its market surveillance authority and adopted draft transposition legislation. Spain built the AESIA, a dedicated supervisory agency, from scratch. Ireland deployed fifteen coordinated authorities under a central National AI Office. Those are genuine structural commitments. Eighteen member states have not reached that point. The extension gives them time. Whether they use it is the bet.

 

The Assumption

The entire framework rests on this: that national competent authorities, operating under 27 different legal frameworks, will converge on consistent enforcement before December 2027. The AI Act is a directly applicable regulation. Its enforcement infrastructure is not. The regulation sets the rules uniformly across the bloc. The authorities responsible for applying them have been built at very different speeds, under very different political conditions. That divergence is the risk the extension is buying time to close. There is no public commitment that the time is sufficient.

 

The Sequence

The AI Act entered into force in August 2024. Member states were required to designate their national competent authorities by August 2025. At least twelve missed that deadline. Seven months later, in May 2026, the Council and Parliament agreed to simplify the rules as part of the Digital Omnibus package. On 29 June, the high-risk AI deadline moved. What remains in force on 2 August is a narrower set: general-purpose AI model obligations and transparency requirements for new deployments. The high-risk AI rules, the Act’s original centre of gravity, are no longer in that set. Governance was adjusted to fit the readiness gap. That is not the order in which enforcement architecture is supposed to be built.

 

The Pager

Lucilla Sioli, Director of the EU AI Office, carries accountability for general-purpose AI enforcement from 2 August. For high-risk AI systems, including credit-scoring models, recruitment tools, and systems used in border control, healthcare, and law enforcement, accountability rests with national competent authorities. In 17 of 27 member states, no public designation exists. The Act names the category. Seventeen member states have yet to name the person.

 

The Proof

The measure that would settle this in 2028 is year-one enforcement consistency: the share of member states that have conducted at least one formal high-risk AI enforcement action, under the same evidentiary standard, in the first twelve months after the December 2027 deadline. No EU institution has publicly committed to publishing that figure. The AI Office’s annual progress reporting is the closest mechanism on the public record. It tracks activity. No published mechanism commits to measuring whether enforcement actions are consistent across member states.

 

Verdict

If the Commission designates a public accountability owner in each member state before December 2026 and commits to publishing year-one enforcement data by name, the 16-month extension holds up as a governance decision made under realistic conditions. Without that, a framework that took two years to reach enforcement hands itself an extension with nobody carrying it.

Plans Don’t Deliver Outcomes. Decisions Do.

The biggest myth in project management is not that it is only about schedules and budgets. That myth was debunked so long ago it barely warrants a mention.

The real myth is more dangerous: that a good plan delivers an outcome.

It does not.

A plan is the document everyone agrees on before the work starts. Delivery is determined by the thousand decisions that happen when that plan meets reality.

 

What a Plan Actually Is

A project plan is a structured expression of intent. It represents the best thinking of a group of people, at a specific point in time, about how they expect work to unfold.

The moment work starts, the plan begins diverging from reality. Not because the planning was poor. Because work is complex, environments shift, and the future is not fully knowable in advance.

The plan does not respond to those divergences. People do.

Someone decides what gets prioritised when two workstreams compete for the same resource. Someone decides what gets descoped when the timeline compresses. Someone decides what gets told to the sponsor and what gets managed quietly at team level. Someone decides whether to hold to the original scope or absorb a late change request that no one has formally costed.

These are not project management artefacts. They are leadership decisions. They happen every day, in every programme, at every level, and the cumulative quality of those decisions determines the outcome, not the quality of the plan that preceded them.

 

What the Data Shows About Plans and Outcomes

McKinsey’s research with Oxford’s Global Projects programme, originally published in 2012 and still McKinsey’s standing figure on its current insights page, based on more than 5,400 IT projects, found that just one in every 200 large IT projects meets all three basic measures of success: on time, on budget, and delivering intended benefits. The same research found that 17 per cent of large IT projects go so badly they threaten the very existence of the company delivering them. Bain’s January 2026 research on reorganisations, based on a survey of nearly 1,000 global executives and employees, found that 88 per cent of company leaders believe their new organisational structure will achieve its goals. Only 36 per cent of the employees actually working inside those structures agree.

These are organisations with project plans. Most of them had quite detailed ones.

The plan was not the variable that determined whether the transformation succeeded. The decisions made inside the transformation were.

McKinsey has been explicit on this, in its analysis of large technology programme management: traditional project management is not built for the complexity of managing a large number of interdependent workstreams. What that observation is really describing is a decision-making capacity problem, not a planning methodology problem.

When multiple workstreams intersect, when dependencies conflict, when assumptions that underpinned the plan prove false, the organisation needs fast, well-informed, appropriately escalated decisions. The project plan cannot make those decisions. A governance structure can enable them, but only if the people inside it are willing and able to act.

 

The Organisations That Deliver

I have worked across a wide range of organisations and programmes. The ones that consistently deliver are not the ones with the most sophisticated planning tools or the most comprehensive project documentation.

They are the ones with a leadership culture that makes fast, honest decisions when the plan diverges from reality.

That culture has specific characteristics. Issues get escalated without penalty. Status reporting reflects what is actually happening, not what the sponsor wants to hear. Scope changes get properly evaluated and decided, rather than quietly absorbed and then discovered six months later as the reason for a cost overrun.

Decisions about resources, priorities, scope, and timing get made by the right people at the right level, at the point when the decision matters, not deferred until the situation has become a crisis requiring emergency intervention.

This is not about removing the plan. A plan is genuinely useful. It creates shared understanding, allocates resources, sequences work, and provides a baseline against which reality can be measured. All of that matters.

But the plan is the starting point, not the delivery mechanism.

 

The Governance Gap Nobody Names

Most programme governance is designed to review progress against plan. Status reports, RAG ratings, milestone trackers, action logs. These are retrospective instruments. They tell you where you have been relative to where you intended to be.

They do not, by themselves, generate decisions.

A programme with robust governance can still fail because the governance structure reports on problems without resolving them. The issues log fills up. The risk register grows. The steering committee meetings run to time, and the programme slides, week by week, toward a late and over-budget delivery, or a cancellation that could have been a scope-reduced success.

The missing element is decision velocity, the willingness and authority to make the calls that change the trajectory, rather than the calls that record that the trajectory has changed.

 

What Good Actually Looks Like

The shift required is not from planning to improvisation. It is from planning-as-delivery to planning-as-baseline.

Build the plan. Use it. Measure against it. But invest as heavily in decision-making culture as in planning rigour. Who has authority to make what decision at what level? How fast can an escalation reach someone with genuine authority? What happens to the person who brings a difficult problem to the steering committee: are they received as someone providing valuable intelligence, or treated as someone who has failed to manage their workstream?

The organisations with the best project outcomes have thought hard about these questions. They are not the ones with the best plans.

They are the ones that can make the right call at 9am on a Tuesday when the plan says one thing and reality says another.

That capacity is the real delivery engine.

Handling Stakeholder Expectations in Digital Transformation: The Honesty Problem

 

Most failed digital transformations were not derailed by technology.

They were derailed by a gap between what was promised at the start and what was achievable in reality. That gap existed from day one, embedded in the business case, and neither the sponsors nor the delivery team chose to address it directly until the programme was already in trouble.

The stakeholder expectation problem in digital transformation is routinely framed as a communication challenge. Better updates. More frequent steering committee engagement. Clearer reporting. These are sensible practices. They are also, in most cases, insufficient, because the problem is not that stakeholders were not kept informed. It is that they were kept informed using numbers and timelines that were not honest about the uncertainty behind them.

That is an honesty problem, not a communication problem. And the fix for it happens at the beginning, not during delivery.

 

The Business Case That Everyone Signed Off On

Digital transformation business cases are almost universally optimistic. Not because the people who write them are dishonest, but because the incentive structure in most organisations rewards ambition and penalises conservatism. A realistic business case, one that acknowledges uncertainty ranges, models downside scenarios, and commits to fewer benefits with higher confidence, is harder to get approved than an ambitious one. So the ambitious one gets written.

The consequence is that the business case becomes a set of commitments rather than a set of hypotheses. By the time the programme moves into delivery, the numbers in the original document are treated as targets rather than as estimates, even when the assumptions underlying them have already been revised. The expectation gap was always there. It was just papered over.

BCG’s analysis of more than 850 companies found that only 35% reach their stated digital transformation goals. Gartner’s October 2024 survey of more than 3,100 CIOs found only 48% of digital initiatives meet or exceed their business outcome targets. The gap is not primarily a delivery capability problem. It is a framing problem.

 

Scope at Altitude

The second structural cause of expectation gaps is scope defined at too high a level of abstraction. Transformation programmes are typically scoped during a phase when the delivery architecture is not yet understood, which means scope boundaries are drawn based on intent rather than on a detailed model of what delivery will actually require.

Both sides, sponsor and delivery team, leave the scoping phase with genuine but different understandings of what is included. Neither party is misrepresenting anything. They simply have not gone deep enough to discover the ambiguity. That ambiguity is then carried into the contract, into the programme plan, and eventually into the steering committee deck, where it will surface as a scope dispute at the moment least convenient for everyone involved.

The fix is not to define scope more tightly in the abstract. It is to define scope at a level of specificity that forces the ambiguity into the open before commitments are made. That is harder and slower than moving quickly to contract. It is also significantly less expensive than managing the dispute six months into delivery.

 

What the Programme Board Actually Hears

Stakeholder management, in practice, often means giving senior sponsors the confidence to remain supportive rather than giving them the information they need to make good decisions. Status reporting in large programmes tends to converge toward reassurance. RAG ratings stay amber longer than conditions warrant, because the consequences of going red feel disproportionate in the moment. Risks that have materialised are carried as risks rather than re-classified as issues. Forecasts are revised gradually rather than reset to reflect the actual picture.

This is not cynical. It is human. Nobody wants to be the person who delivers bad news. The programme team has worked hard. The delays feel temporary. There is always a reasonable argument for holding the line a little longer.

The problem is that by the time the gap between expectation and reality is reported honestly, it is too large to close without a significant reset. The reset conversation is much harder than it needed to be, because the sponsor was not kept informed of how the gap was developing.

 

The Expectation Reset Conversation

When the gap surfaces, and it always surfaces, the response that preserves the programme is not to defend the original business case. It is to reframe the conversation around what is still achievable, with what degree of confidence, on what timeline. That requires the delivery team to be willing to put a revised view in front of the sponsor, acknowledge that the original framing was overoptimistic, and propose a credible path forward.

That conversation is significantly easier when the relationship between sponsor and delivery has been built on honest reporting from the start. It is significantly harder when the sponsor has been receiving optimistic status updates and now feels misled, not because anyone intended to mislead them, but because the communication was shaped by the desire to maintain confidence rather than the obligation to maintain accuracy.

 

Less Ambiguity, Earlier

The conditions that produce the stakeholder expectation problem are well understood: optimistic business cases, scope defined at altitude, and status reporting shaped by the incentive to maintain confidence. None of these are inevitable.

The organisations that manage stakeholder expectations well are not the ones with the best communication strategies. They are the ones with the discipline to be specific about uncertainty before programmes begin, to define scope at a level of detail that surfaces ambiguity early, and to build reporting practices that give senior sponsors the information they need to make real decisions rather than the reassurance they want to maintain support.

Less ambiguity earlier. Fewer numbers presented as certain when they are estimated. Fewer commitments made before the delivery architecture is understood.

That is the fix. It is harder to sell at the outset and significantly easier to live with throughout delivery.

Pre-Mortem: Eight Companies, No Published Accountability Standard

The Pre-Mortem is a weekly series on this blog. Each piece applies five questions to a major technology commitment before the outcome is known.

In February 2026, the United States Department of War signed agreements with eight of the world’s leading artificial intelligence companies, OpenAI, Google, Microsoft, SpaceX, Oracle, Amazon Web Services, NVIDIA, and Reflection, to deploy their advanced AI models inside its classified networks. Impact Level 6 (IL6) covers data classified at the Secret level. Impact Level 7 (IL7) covers compartmented intelligence and the most sensitive operational systems, where the United States military runs its actual warfighting decision support. This is the first time that large language models have operated within IL7 environments. What has not been published is who carries accountability when one of them gets something wrong.

 

The Bet

The Department of War’s stated aim is to establish the United States military as an AI-first fighting force, achieving what its AI Acceleration Strategy calls decision superiority across all domains of warfare. The eight agreements are the mechanism. The AI systems will summarise surveillance feeds, synthesise intelligence data, and suggest tactical options to human operators. The Department of War’s five AI ethics principles, responsible, equitable, traceable, reliable, and governable, are on the record. The bet is that those principles are sufficient architecture for what happens inside a classified environment.

 

The Assumption

The whole bet turns on this: that “humans remain accountable for AI outcomes” as a stated principle is equivalent to a published accountability framework.

That distinction is where there is a gap. The Department of War’s Responsible AI Strategy and Implementation Pathway establishes process. It does not name the specific individual, command role, or governance layer accountable when an AI-assisted intelligence summary inside an IL7 environment shapes a decision that turns out to be wrong. Principle and framework are not the same thing, and in a classified environment that distinction cannot be tested publicly.

 

The Sequence

In July 2025, Anthropic’s Claude became the first frontier AI model approved for use on classified networks. The Pentagon subsequently sought to renegotiate those terms, demanding Anthropic permit its models to be used for all lawful purposes without limitation. Anthropic declined, citing concerns about mass domestic surveillance and autonomous weapons. On 27 February 2026, President Trump ordered all federal agencies to stop using Anthropic. The following day, OpenAI signed its classified deal with commitments that included prohibitions on domestic mass surveillance and human responsibility for the use of force, positions that aligned with the guardrails Anthropic had sought to retain. By May 2026, the remaining seven of the eight, Google, Microsoft, SpaceX, Oracle, Amazon Web Services, NVIDIA, and Reflection, had signed equivalent agreements.

The sequence reveals something structural. The accountability architecture for classified military AI was settled by commercial negotiation and political designation, not by a published governance framework.

 

The Pager

Legal scholars on autonomous weapons identify the same accountability fracture that applies in the decision-support context here. When an AI-assisted output causes harm in a classified environment, accountability distributes: software developers could not have anticipated all operational contexts, commanding officers disclaim responsibility for machine-generated outputs, vendors invoke contractual limitation of liability. The human-in-the-loop design means a person reviews AI suggestions before acting. It does not mean accountability for acting on a wrong AI output has been named anywhere in the command chain.

No published document names the specific individual role, command layer, or governance body accountable for a wrong AI-assisted output inside an IL7 environment. No congressional oversight mechanism covers classified operational AI use. No published error reporting standard exists. By the nature of classified operations, none can.

 

The Proof

Eight companies, the highest classification levels, large language models operating on top-secret data for the first time: the scale of the commitment is confirmed. The outcome data will not follow. Classified operational AI performance is not publicly reviewed, by design. This is the only deployment in this series where the proof question cannot be answered from the outside, not because the data is not collected, but because it cannot be published.

The accountability question is not whether humans are in the loop. They are, by stated commitment. The question is whether the framework for who carries it specifically, when they get something wrong, inside a system that cannot publish what it got wrong, exists in any enforceable form.

 

The Verdict

If the Department of War’s five principles are operationalised into a named, enforceable command accountability chain for AI-assisted decisions at every classification level, if the commercial guardrails in all eight agreements are independently verifiable by a body with appropriate clearance, and if a congressional oversight mechanism specific to classified AI operational failure is established, then this is what responsible military AI deployment at scale should look like.

Without all three, eight of the most powerful AI systems on earth are running inside the most classified networks in the world. The decisions they shape will not be publicly reviewed. The wrong ones will not be counted.

The accountability is a principle. The framework has not been built yet.

AI Gets the Blame. Governance Built the Problem

When an AI system causes serious harm, the story is easy to write. Biased algorithm. Dangerous model. Technology that cannot be trusted. The headline names the AI and moves on.

What the headline does not name is the governance meeting that never happened. The audit trail that was never built. The accountability structure that should have existed before the first automated decision was made, and did not.

That absence is the story.

Across healthcare, hiring, public services, and criminal justice, the pattern repeats with such consistency that it has stopped looking like bad luck and started looking like a structural failure. AI is deployed to solve a genuine operational problem. The deployment decision is made without the governance architecture that would constrain it, audit it, or catch its errors before they compound. Consequential harm compounds. And then AI gets the blame, while the process failure is buried somewhere in paragraph twelve.

These are not AI stories. They are governance stories. AI is the mechanism.

 

The Cases That Make It Visible

In January 2026, a class action was brought against Eightfold AI, a hiring platform used by major employers globally. The case was filed by Jenny Yang, former chair of the Equal Employment Opportunity Commission. It does not argue that the algorithm was biased, though that question remains open. It argues that the system operated in secret. Eightfold had scored over one billion workers on a scale of zero to five, and candidates ranked at the bottom were discarded before a human being ever saw their application.

Seventy per cent of companies using AI in hiring allow AI to reject candidates at the initial screening stage, with no human review at that point. One in five goes further, allowing AI to reject candidates at every stage of the process with zero human involvement at any point. That is not a technology decision. It is a governance decision. Someone, somewhere, made the deliberate choice to remove the human from the loop. Nobody built an accountability structure around what happened next.

The pattern is not new. In 2021, the Dutch government resigned after an AI system falsely accused twenty thousand families of child welfare fraud. Courts ordered repayments of tens of thousands of euros per family. In Australia, the Robodebt programme issued four hundred thousand wrongful fraud accusations before it was ruled unlawful and the government repaid over one billion dollars. In Michigan, a 2024 settlement reimbursed three thousand plaintiffs for what a benefits fraud algorithm had wrongly taken from them.

In each case, the AI system did what it was designed to do. What nobody designed was the mechanism that would question whether it was doing the right thing.

 

The Structural Argument

The research makes the pattern numerical.

An analysis of a hundred and forty enterprise AI implementations found that only twenty-three per cent of failures were caused by model performance, data quality, or technical integration. The remaining seventy-seven per cent came down to strategy, governance, and change management.

Three-quarters of AI failures have nothing to do with the technology.

Only one in five organisations has a mature governance model for autonomous AI agents. This is the population deploying AI at scale, across consequential decisions in hiring, healthcare, benefits, credit, and criminal justice, mostly without the mechanisms needed to know whether the AI is producing correct outcomes, or what to do when it does not.

This is not a portrait of reckless technology. It is a portrait of reckless deployment. The AI worked. The organisation around it did not.

 

Governance as Delivery Discipline

Most organisations treating AI governance as a compliance function are building the next wave of failures right now. Compliance asks whether the system meets a threshold at the point of deployment. Delivery discipline asks whether the system is behaving as intended across every subsequent decision it makes, and whether there is anyone accountable when it does not.

These are not the same question. That gap is where accountability ends and harm begins.

Effective AI governance is not about slowing deployment. It is about building the accountability architecture alongside the deployment. An agreed definition of what the system is supposed to achieve. A method for measuring whether it is achieving it. A human being, with authority and accountability, responsible for reviewing outcomes at meaningful intervals. An appeals mechanism when the system gets it wrong, because it will. Documentation that allows an audit when something goes wrong, rather than after the damage is done.

 

The Question Worth Asking Before the Next Deployment

The Eightfold case will not be the last of its kind. The healthcare billing figures will grow before they shrink. More governments will face the political and financial cost of systems that automated consequential decisions without the mechanisms to catch errors before they multiply.

The organisations that avoid this are not the ones that move slower on AI. They are the ones that treat governance as part of what delivery means, not as a separate conversation to have later, when the headlines arrive.

By then, the structure has already failed. The question worth asking now is whether yours is being built.

Your Enterprise AI Programme Is Structured Backwards.

There is a research paper that has been making rounds in enterprise AI circles and deserves more attention than the single line most people have taken from it.

The Stanford Digital Economy Lab published its Enterprise AI Playbook in April 2026. It is drawn from 51 live, production-grade deployments across 41 organisations, seven countries, and more than one million employees. And what it found cuts against almost every assumption the technology industry has built its messaging around.

In every deployment that succeeded, and in every deployment that failed, the determining factor was not the model. It was not the vendor, the feature set, the benchmark performance, or the integration timeline. In every case, the variable that decided outcome was the organisation: executive sponsorship, governance architecture, and the quality of workforce change management. Seventy-seven per cent of the implementation challenges in the study traced to non-technical factors: change management, data quality, and process redesign.

 

One Finding in 51 Deployments

The research does not say technology does not matter. It says it matters significantly less than most programmes treat it as mattering, and that the organisations which led with governance and change management consistently outperformed those that led with model selection.

One adjacent data point makes this more concrete. Sixty-one per cent of the successful deployments in the study followed at least one prior failed attempt. Those organisations had not found a better model on the second attempt. They had changed what they were actually doing. And in almost every case, that meant addressing the organisational variables they had underestimated the first time: governance structure, change management approach, and clarity of executive ownership.

The technology was not the lesson. The organisation was.

 

How Most Programmes Are Actually Structured

I have sat in enough enterprise AI programme kick-offs to recognise the pattern before the second slide.

The programme begins with a vendor selection process. A proof of concept is scoped, model performance is evaluated, pricing tiers are compared, latency is benchmarked. The technology conversation consumes the first two to six months of the programme. By the time it concludes, the organisation has committed significant capital and credibility to a specific platform before the questions the Stanford data confirms are the real determinants of success have been seriously engaged.

Those questions are not complicated. Who is the executive sponsor, and what does their sponsorship mean in terms of decision-making authority and resource commitment, not just endorsement? What is the governance architecture for the AI programme, not for the AI system, but for the programme? How does the organisation plan to manage the workforce transition that a serious deployment requires, and what does it know about the change-readiness of the teams it is deploying into?

These are not questions that get answered in a vendor evaluation. They are not questions that appear on most programme charters. They are the questions that decide whether the programme succeeds.

 

What Leading With Governance Actually Means

Leading with governance does not mean delaying deployment while a committee produces documentation nobody will read. It means defining, before the technology is in the ground, who owns the programme’s outcomes, how the workforce transition will be handled, and what the decision-making structure looks like when the deployment hits the friction that every serious AI implementation hits at scale.

Because the friction will come. It always does. And how an organisation responds to it reveals which frame it used at the start.

Programmes built on a technology frame diagnose the friction as a technology problem. The model is adjusted. The integration is patched. The interface is redesigned. The organisational dynamics actually driving the resistance go unexamined, because the programme was never looking at them. Programmes built on a change management frame diagnose the same friction differently. The conversation shifts to whether the right people were involved in design, whether transition support was adequate, whether the governance gave teams the clarity they needed to work confidently with the new system. Those questions lead somewhere. The technology-first version usually leads to another vendor call.

 

The Argument That Is Now Evidence

This is not a new insight for experienced transformation leaders. I have been making a version of this argument for years, and so has almost everyone else who has led a serious enterprise change programme. The frustration, and the genuine value of Stanford’s work, is that it can now be asserted with data.

For transformation leaders making the case for governance investment in leadership conversations where the pressure is almost always to accelerate on technology, the Stanford Playbook is the data point that turns an argument from opinion into evidence. It does not require arguing against the technology. It requires arguing about sequence and proportion, and it gives you the empirical foundation to do it.

The organisations still leading with model selection are systematically delaying the decisions that actually determine whether a deployment succeeds.

Fifty-one real deployments confirm it. That should be enough.

Your AI Risk Register Does Not Reflect Your Actual Risk

 

On 22 June 2026, the intelligence agencies of the United States, United Kingdom, Australia, Canada, and New Zealand spoke in a single voice about enterprise AI risk, and what they said demands attention.

The Five Eyes cybersecurity agencies issued a joint statement warning that frontier AI models are improving at a pace that will allow them to bypass prevailing enterprise cybersecurity defences within months. Not within years. Not in the next planning cycle. Within months. The statement’s own language: “The timeline is not years, it is months.”

 

This Is Not an Abstract Warning

Joint statements from the Five Eyes agencies carry a different category of authority than vendor advisories or consultancy threat reports. These are national intelligence services with access to classified threat intelligence, speaking to government and enterprise leaders simultaneously. When they frame a risk as both imminent and enterprise-specific, take it at face value.

What sets this advisory apart from every AI security conversation most enterprises have been having is one thing: specificity. The Five Eyes statement does not describe abstract AI risks. It specifically names the enterprise AI tools deployed at scale in the last 18 months: copilots, AI assistants, browser-connected agents, and systems with access to operational and customer data. The primary attack mechanism, developed across Five Eyes guidance published earlier this year, is prompt injection: an adversary embeds hidden instructions in content the AI system processes, causing it to act outside its intended scope.

That specificity matters. It means the tools that most large enterprises have already deployed are the attack surface being described.

 

The Threat Moved Faster Than Your Review

Most organisations that have rolled out AI copilots, enterprise agents, or browser-integrated assistants have conducted security reviews of those deployments. The Five Eyes advisory is not questioning whether those reviews happened. It is saying that the threat has moved faster than the defences, and that a review conducted six months ago may no longer accurately reflect the risk profile today. The gap is not in intent. It is in elapsed time against a threat that has not stood still.

The advisory is explicit that this is not solely a security-team problem. The statement directs its recommendations at leadership, framing AI-driven cyber risk as a governance and board-level accountability question. The statement’s own title: “The AI shift in cyber risk: why leaders must act now.” That framing has direct implications for how risk registers are built and how AI deployment decisions are reported to boards.

 

Three Things Worth Doing Before Your Next Board Meeting

The advisory points to three things transformation leaders should act on before their next board meeting.

The first is a current security review. Every AI deployment connected to operational data, whether customer records, financial systems, or internal communications, needs a review that specifically addresses prompt injection risk. Not the review conducted at go-live. A current one, calibrated to the threat capability the Five Eyes describe as arriving within months.

The second is an updated risk register. Most enterprise risk frameworks assessed AI security risk at the point of initial deployment. The Five Eyes advisory says the threat environment has changed materially in the months since, and the assessment needs to reflect current threat capability rather than historical assumptions. An outdated risk assessment is not a minor administrative gap at this point. It is a governance exposure.

The third is using the advisory to reframe the conversation at board level. Six cybersecurity agencies from five countries issued this statement with an explicit focus on business leadership. That gives transformation leaders the instrument they need to move boards that have been treating AI security as an implementation detail. The Five Eyes advisory makes it a governance question. Use it as one.

The AI deployment decisions taken in the last 18 months created an attack surface. Most enterprise risk registers have not yet priced what that surface is worth to an adversary with AI-powered attack tools that are months from bypassing prevailing defences. That gap needs to close, and it closes with a current assessment, not one accurate at the time of go-live.