93% Have Been Breached by Vulnerable AI Code. 30% Still Ship It Anyway.


Seventy per cent of developers believe AI-generated code carries more vulnerabilities than the code they write themselves. Thirty per cent ship it into production anyway, knowingly. Ninety-three per cent of the same respondents report at least one breach traced back to a vulnerable application. Awareness of the risk and action on the risk are not the same thing. They rarely are, according to every major AI risk study published so far this year.

 

The Core Finding Repeats Across Every Study That Looks

The Purple Book Community’s State of AI Risk Management 2026 report, surveying more than 650 senior cybersecurity leaders across North America and Europe between December 2025 and February 2026, found 59 per cent of organisations acknowledging shadow AI within their own environment, despite 90 per cent claiming confidence in their AI visibility. Seventy-eight per cent are already piloting or deploying agentic AI systems, and 73 per cent say AI-assisted development is outpacing their security review cycles. Fifty-one per cent run 11 or more separate security scanning tools, and 46 per cent of teams spend significant time triaging vulnerabilities that turn out not to matter. Governance has not kept pace with adoption. It has fallen further behind with each new AI capability organisations switch on.

 

The Global Maturity Picture Is Worse Than Any Single Region’s

The World Economic Forum’s Advancing Responsible AI Innovation research, covering 1,500 organisations worldwide, found 81 per cent still sitting in the earliest two stages of responsible AI maturity. Regional breakdowns make that global figure look almost optimistic. In Asia-Pacific, only 1 per cent of organisations have fully operationalised responsible AI practices, a gap researchers describe as more pronounced than the global average. Separate research from Dataiku found 94 per cent of CEOs suspect employees are already using generative AI tools without authorisation, while 75 per cent of data leaders admit low trust in their own organisation’s AI agent deployments.

 

The Middle East Shows a Different Symptom of the Same Disease

Middle Eastern enterprises are not struggling with pilots. Multiple regional AI advisory reports this year describe successful proof-of-concept projects stalling the moment organisations attempt enterprise-wide rollout, blocked by exactly the governance gap the global data points to: no clear ownership of AI risk decisions, inconsistent access controls, and the added complexity of navigating different data protection and AI-specific regulations across Gulf jurisdictions simultaneously. The technology performs in the pilot. The governance structure needed to scale it safely usually does not exist yet.

 

What This Means for How Organisations Govern AI Risk

Every regional variant of this research points to the same structural gap rather than three unrelated problems. A risk committee should treat confidence surveys as a warning sign rather than reassurance, ask specifically what percentage of AI-generated code has been reviewed rather than whether a review process exists on paper, and confirm who owns AI risk decisions before the next agentic AI pilot moves toward production. The pattern holding across North America, Europe, Asia-Pacific and the Gulf is consistent enough to stop treating it as one company’s oversight and start treating it as this year’s defining governance failure.

The Difference Between an AI Pilot and an AI Programme Is Who Gets Blamed When It Fails.

A pilot has no name attached to its failure. A programme does. That is the entire distinction, and almost nobody treats it as the one that matters.

Ask most organisations why their AI initiative is still called a pilot eighteen months after launch and the answer usually involves budget, integration complexity or waiting for the model to mature. The real answer is simpler and less comfortable. Nobody has agreed who owns the outcome if it goes wrong, and a pilot is the one structure where that question never has to be answered.

 

The Pilot Was Never Built to Answer This Question

MIT’s Project NANDA found that 95% of generative AI pilots fail to deliver a measurable return, based on 150 leadership interviews, a survey of 350 employees and analysis of 300 public deployments. The report itself points to a different explanation: tools that never adapt to how the organisation actually works, budgets aimed at sales and marketing while the real return sits in back-office automation, and internal builds that consistently underperform specialist vendor partnerships.

Underneath all three is the same gap. A pilot is built to prove a capability exists. Whether anyone is answerable for what happens once that capability touches real customers, real decisions and real money is a different question, one a pilot was never built to answer. Most organisations discover this the hard way, months into a pilot that technically works and still cannot get budget to go further, because nobody signed up to own what happens next.

 

Ownership Is the Line, Not Scale or Budget

Grant Thornton’s 2026 AI Impact Survey of 950 senior leaders found 78% lack strong confidence they could pass an independent AI governance audit within 90 days. Among organisations still piloting, that confidence drops further, to just 7%. The gap shows up directly in results: organisations with fully integrated AI are close to four times more likely to report AI-driven revenue growth than those still piloting, 58% against 15%.

Grant Thornton’s Tom Puthiyamadam put the underlying issue plainly: organisations that have invested in governance move faster precisely because they have the confidence to scale, while the ones without it are one incident away from a far harder conversation.

MIT’s own findings point the same way: among the factors separating pilots that scale from the ones that stall, the report names empowering line managers, not just centralised AI labs, to drive adoption. A central lab can build a working model, but naming who answers for what that model does inside someone else’s process is a separate task entirely, one a programme takes on and a pilot leaves undone.

 

What a Programme Actually Commits To

Turning a pilot into a programme is not a budget decision. It is naming, before the next phase starts, who is accountable if the thing fails in production, what failing actually means in that specific context, and what happens in the following week if it does. None of that requires new technology. It requires a decision most organisations postpone precisely because a pilot lets them.

Just 38% of organisations have a formal, comprehensive AI policy in place, up from 28% the year before, according to ISACA’s 2026 research covering more than 3,400 digital trust professionals globally, which means most of what currently passes for AI governance gets improvised the first time something breaks, rather than designed before it ships. An owner named after an incident is not accountability. It is damage control wearing accountability’s name.

The organisations closing the gap treat this as day-one work, not late-stage paperwork. A named business owner, not a technical one, accountable for the outcome. And a clear definition of what failure looks like for that specific use case, agreed before launch rather than improvised during the post-mortem.

 

Whose Name Is on This When It Breaks

Before the next AI initiative gets called a programme instead of a pilot, ask one thing in the room where budget gets approved: if this fails next month, whose name is on the outcome, and did they agree to that before it happened or only after.

Most organisations cannot answer that today. That gap explains why so many pilots never leave the lab.

Governance Fatigue Is Real, and It’s Killing the Governance That Matters

Every governance failure gets the same response: add a committee. Nobody ever asks which of the five committees already in the room should be deleted.

I have watched this play out on the same programme, twice, thirteen months apart. An incident happens. A review is commissioned. The review recommends a new gate, a new sign-off, a new board with a name that sounds important. Nobody asks whether the five existing boards had already covered this ground and simply were not being used properly. The new layer gets built. The old layers stay exactly where they were, because retiring a control is a much harder conversation than adding one, and nobody wants to be the person who removed the safeguard right before something went wrong.

Multiply that pattern across a few years of incidents, mergers, audits and regulatory nudges, and you get an organisation with more governance than anyone can actually operate.

 

The sprawl nobody planned and everybody built

Governance frameworks get bloated one reasonable-sounding addition at a time, not in a single decision. A near-miss produces a new checkpoint. An audit finding produces a new form. A departing executive leaves behind a committee that made sense under their sponsorship and none under anyone else’s. Each addition was defensible in isolation. Nobody ever sat down and asked what the whole structure looked like once you added them all together.

The organisations most proud of their governance maturity are often the ones carrying the heaviest version of this problem. More boards. More gates. More documented sign-offs. It reads as rigour on an org chart and feels like wading through treacle to anyone actually trying to get something delivered.

 

Fatigue looks like silence, not rebellion

The expensive part is what people do once the meetings stop making sense to them, not the extra meetings. They do not object. They do not escalate the absurdity of an ninth sign-off. They quietly learn which boxes can be ticked without real scrutiny, which approvals are theatre, and which route gets something through fastest regardless of whether it is the correct one.

That is governance fatigue, and it is far more dangerous than having too little governance in the first place. An organisation with no controls at least knows it is exposed. An organisation with too many controls believes it is protected, right up until the one decision that actually mattered slipped through a gate everyone had stopped taking seriously.

 

Clarity beats volume, every time

The research on decision rights backs this up more directly than most governance debates acknowledge. Itonics’s analysis of partner and programme governance found that organisations using a properly maintained RACI framework report 70 per cent fewer “who decides” disputes and 25 per cent faster decision cycle times. RACI works by replacing ambiguity with a single, shared answer to a question that used to require a meeting to resolve, not by adding another layer. The gain comes from governance that is unambiguous enough that people stop needing to ask, rather than from adding more of it.

That is the distinction most organisations miss when they respond to a failure by adding structure. The problem was usually oversight so diffuse that nobody could say, without checking three separate documents, who actually held the decision.

 

The discipline of taking something away

Fixing this requires a habit most organisations have never built: retiring governance on purpose. Every new control should come with an audit of what it is replacing, not just what it is adding. Every steering board should have to justify its existence against a simple test: if this group disappeared tomorrow, what decision would genuinely not get made anywhere else? If the answer is “nothing, it would just move up a level,” that board is inertia with a calendar invite, not governance.

The organisations that manage this well treat their governance structure the way a good engineer treats a system under load, asking what is carrying weight it no longer needs to carry and taking it out, rather than just adding capacity when something breaks.

Governance was never supposed to be heavy. It was supposed to be clear. Somewhere along the way, most organisations mistook the two for the same thing.

 

What We Found When We Measured Adoption, Not Just Deployment.

A go live report is easy to write. Every site is on the new system, every licence is issued, every training session delivered on schedule. Three months later, the usage dashboard tells a different story, and it is the dashboard nobody puts in front of the steering committee.

 

The Metric Everyone Reports

Deployment is countable in a way adoption never is. Percentage of sites migrated, number of licences activated, hours of training delivered, these are the figures that go into a programme status report because they can be measured on the day the rollout finishes. None of them says whether anyone is still using the system a quarter later, or whether they have quietly gone back to the spreadsheet it was meant to replace.

 

The Number Nobody Puts in the Steering Deck

IBM’s 2026 Global CEO Study, based on more than 2,000 chief executives surveyed worldwide by the IBM Institute for Business Value, found that only 25 per cent of workers use AI regularly in their jobs, even though 86 per cent of CEOs believe their workforce already has the skills to do so. Eighty three per cent of the same CEOs said AI’s success depends more on people’s adoption than on the technology itself. The rollout finished on schedule. The adoption did not follow.

 

What Poor Adoption Actually Costs

A Forrester Consulting study commissioned by Whatfix, surveying 335 senior decision makers at large organisations across North America, Europe, APAC and India, put a figure on what that gap costs a mid-sized enterprise, $10.9 million a year, plus 728 hours lost per employee navigating systems that were rolled out but never properly embedded. That is not a training budget line. It is the ongoing cost of a deployment nobody followed up on. The same research found a wide gap between organisations with adoption maturity and those without, 53 per cent of mature adopters reported improved user experience against 28 per cent of the rest, and 56 per cent reported stronger return on investment against 28 per cent.

 

Why Deployment Metrics Miss This

Programme reporting is built around milestones a PMO can close off: go live achieved, training complete, licences distributed. Adoption behaves more like a curve than a milestone, one that keeps moving long after the project has been marked complete and the team has moved on to the next initiative. By the time low usage shows up in a satisfaction survey or a renewal conversation, the people who owned the rollout are three programmes further down the roadmap.

Part of the reason adoption rarely gets measured is that almost nothing in a typical programme is set up to reward it. Vendor contracts are frequently structured around go live milestones rather than usage thresholds, so the commercial incentive to keep measuring stops the day the system switches on. Programme teams are resourced to deliver a rollout, not to own what happens to it afterwards, and by the time adoption data would be available, the team has usually been reallocated to the next initiative. None of this is deliberate. It simply reflects a reporting structure and an ownership structure that both end at the same milestone.

 

What We Started Measuring Instead

On some programmes I have run, the fix had less to do with better software and more to do with what the steering committee agreed to look at. We started tracking active usage at 30, 60 and 90 days after go live, alongside a simple drop off rate, how many people who logged in during week one had stopped logging in by week twelve. We also moved benefits realisation sign off away from the go live date and tied it to a usage threshold instead, so a project could not be closed as successful until people were actually using what had been built. It is a small governance change, and it surfaces problems a deployment report never will.

Rollout dates and licence counts still matter, deployment discipline was never the issue. What most programmes are missing is a second dashboard sitting next to the first one, what deployment made possible, and what adoption is showing three months after anyone stopped watching. The programmes that quietly fail are rarely the ones that missed a go live date. They are the ones that hit it, reported it as a win, and never checked what happened next.

Your Gates Aren’t Protecting the Business. They’re Protecting Themselves.

Nobody sets out to build a bureaucracy. Every heavy stage-gate process started as three good intentions: get bad projects killed early, get good projects through fast, and keep a clean record of why each call was made. Then it grew a review board nobody remembers approving, and the good projects started waiting as long as the bad ones.

That is the actual failure. Not that gates exist. That almost nobody still measures them against the job they were built to do.

 

The Test a Stage Gate Was Actually Built to Pass

Governance run properly delivers three things, and nothing else matters as much as these: faster decisions, so a good project stops waiting weeks for a yes. Earlier kills, so a weak one frees up capital instead of quietly draining it for another two quarters. And a clean audit trail, so nobody has to reconstruct the reasoning after the fact. Get those three right and a stage gate speeds decisions. Miss them and it slows every decision down, good and bad alike, because the mechanism has stopped doing the job it was built for.

Someone genuinely has to decide which projects live, which pivot, and which ones are quietly draining the business. That much was never in question.

 

Why the Mechanism Rots

Four patterns do most of the damage, and each one accumulates quietly rather than arriving as a single bad decision. Entry gates get heavy while exit and kill discipline stay weak or disappear entirely, so zombie projects clog the funnel and starve the strong ones of attention. Panels grow larger and meetings grow longer until authority is spread across so many people that nothing actually gets decided. Reviews turn into theatre, rubber-stamping or deferring rather than choosing, and momentum dies in the gap between gates. And decision rights stay ambiguous enough that nobody is the clear owner of the yes or the no, so everything escalates and stalls at once.

Each of those four is a governance design that stopped serving the teams running through it and started serving itself, not a process flaw you fix by adding another step. The entry gate feels productive, so it keeps growing. The exit gate feels harsh, so nobody wants to own it, and that imbalance is where most of the trouble actually starts.

 

The Numbers Behind the Frustration

The frustration shows up in real operating numbers, not just complaints. At Vivix Vidros Planos, a Brazilian flat-glass manufacturer, resolving a customer complaint used to take weeks, long enough for a buyer to lose patience and take the next contract elsewhere. After building an AI-powered chatbot into its existing production and quality data, that shrank to minutes, an 80% reduction in complaint resolution time. Responses to production-line issues sped up by a further 85%. Both figures come from Vivix’s own case study, published jointly by Siemens and AWS. The underlying pattern holds regardless of the exact numbers: the real cost of a slow gate shows up as lost trust and lost contracts, not just lost hours.

 

What Minimum Viable Governance Actually Looks Like

The fix is sizing each gate to the actual risk in front of it, rather than running every decision through the same heavy process regardless of what it actually requires. A low-risk process tweak does not need the same panel as a bet-the-quarter platform launch, and treating both the same is exactly how standing committees fill up with work they should never see in the first place. High-risk decisions keep full board review, because that rigor is proportionate there. Mid-tier decisions get a lightweight gate with a single accountable owner. Low-risk work proceeds by default through a simple intake form, with oversight applied only if something in it actually warrants it.

The useful design target sits between two failure modes: above a ceiling, controls become bottlenecks and teams quietly route around them; below a floor, real risk creeps in unmanaged. Getting that band right is not abstract. One organisation that tightened its policy down to the minimum viable version halved the time complex decisions took and surfaced three times more opportunities than peers still running the heavier version.

 

The Test Most Gates Would Fail

Pick the last three projects your organisation killed at a gate, and the last three it approved. If the kills took longer to reach than the approvals, the gate is not protecting the business from bad decisions. It is protecting itself from having to make any decision at all.

Pre-Mortem: The Accountability Question the Mills Review Left Open

On 6 July 2026, the Financial Conduct Authority published the Mills Review, its examination of how AI will reshape retail financial services in the UK. The review covers seven recommendations across the regulatory perimeter, oversight architecture, and the transition to autonomous decision-making. It names the accountability gap at the centre of autonomous AI trading. It does not close it.

This is the thirteenth piece in the Pre-Mortem series. Five questions, applied to the public record, before the outcome is known.

 

The Bet

UK firms deploying autonomous trading AI are betting that the Senior Managers and Certification Regime (SMCR), the framework that holds named executives personally accountable for conduct failures in their area of responsibility, covers their position through general senior manager oversight. The FCA has been clear that delegating a decision to an algorithm does not transfer senior manager liability to the algorithm. The bet is that this principle, correctly stated and on the public record, can be demonstrated in practice before an enforcement case defines what demonstrating it actually requires.

 

The Assumption

Seven recommendations. One question still without an answer:

When an autonomous trading system executes a decision at machine speed, without pausing for human approval of the individual trade, which specific senior manager function is accountable if that decision causes a customer loss, and what does demonstrating adequate oversight of a system like that actually require?

The Mills Review acknowledged the problem directly. Without guidance, the review found, the combination of greater opacity in AI-mediated decisions and factors such as model drift makes it harder for the regulator to identify a de facto responsible individual, or for senior managers to evidence meaningful human control. Stakeholder feedback throughout the review called for clearer guidance on what constitutes the “reasonable steps” expected of senior managers. The review recommends the FCA develop it. The FCA has not yet published it. Every firm currently deploying autonomous trading AI is operating on the assumption that its existing accountability structure covers the gap. That assumption has not been tested in an enforcement case.

 

The Sequence

December 2019. SMCR extended to all FCA solo-regulated firms, completing its rollout across financial services.

27 January 2026. The FCA launched the Mills Review, acknowledging that AI in retail financial services had developed faster than the regulatory frameworks designed to govern it.

24 February 2026. Call for input closed.

6 July 2026. The review published seven recommendations. The FCA committed to adapting its regulatory frameworks as the transition to autonomous models continues. No guidance named a specific senior manager function as accountable for autonomous trading decisions. No guidance defined what “reasonable steps” requires for a system executing at machine speed without human review of individual decisions.

The capability reached the market before SMCR was tested against it. The review arrived after the capability. The guidance has not arrived yet.

 

The Pager

The FCA has confirmed there will be no dedicated Senior Manager Function for AI, and that accountability falls on existing functions. That is a clear policy position and it deserves credit for being stated plainly. The Treasury Select Committee has urged the FCA to publish guidance specifying the level of assurance expected of senior managers for AI-related harm. The Mills Review carried that request forward into its recommendations. The harder question is the one seven recommendations did not answer: when an autonomous trading system causes a customer loss, which specific function holder carries the call?

 

The Proof

There are no enforcement cases. The first case will establish what “reasonable steps” means in an AI trading context. The Mills Review is a process measure: it produced recommendations. The outcome measure worth watching is whether the FCA’s follow-on guidance names a specific function and defines the oversight standard in operational terms rather than principles alone. A principle restated is not a gap closed.

 

Verdict

If the FCA’s follow-on guidance names the senior manager function accountable for autonomous trading AI and defines what “reasonable steps” requires at the operational level, UK financial services will have resolved an accountability gap that every other major jurisdiction is still navigating. The review’s existence, the named individual who led it, and the seven published recommendations are genuine evidence that the FCA identified the problem and moved on it. Without operational guidance, the gap stays open. The first enforcement case will write the rule in the least comfortable setting available. That is a considerably worse way to write it.

Why Traditional Project Management Is Failing Modern Teams

Why Traditional Project Management Is Failing Modern Teams

Most project failures get blamed on execution. A missed deadline. A stakeholder who went quiet at the wrong moment. A scope that crept until nobody could point to when it happened.

Look earlier and the failure was already built in before a single sprint started.

A Forbes Technology Council analysis makes the point directly: misalignment gets embedded into the foundation long before execution begins, not manufactured somewhere in the middle. The team that won the deal rarely stays involved in delivery. The customer’s actual operating mindset only reveals itself once work is already moving. The incentives written into the contract often point delivery and client in different directions before day one. Teams execute a plan that was already structurally unsound before the first sprint started.

Traditional project management was never built to catch that kind of problem. Waterfall assumes you can define requirements fully upfront, lock them, and deliver against a fixed spec months later. That assumption survives about as long as the first change request. Priorities that shift inside a six-week planning cycle, which describes most programmes now, make a locked spec obsolete before it has even shipped.

 

Rigidity Is the Symptom. Something Else Is the Disease.

Here’s the twist most framework debates miss. Methodology by itself isn’t what separates the teams that deliver from the ones that don’t. PMI’s most recent Pulse of the Profession research found project performance sits at roughly 73.8% whether a team runs predictive, hybrid, or agile delivery, and whether people work remote, hybrid, or in-person. What actually moved the needle was business acumen: professionals strong in it posted 27% lower failure rates, regardless of which framework sat on the wall.

Framework still matters. It was just never the whole story, and treating “which methodology” as the central question misses where most programmes actually break.

A PM Solutions case study makes the same point at a larger scale. A U.S. staffing company with more than 8,000 internal and 90,000 contract employees had already tried and failed to stand up a PMO once. On the second attempt, the team built a hybrid methodology suited to the client’s actual environment, blending traditional project management, agile, and the touchpoints where each meets software delivery, then added a governance structure, portfolio visibility and resource planning on top of it. Within six months, every project flagged red under the new reporting system, more than $13 million worth of work, was recovered. One severely troubled multi-year project, over budget and behind schedule, was turned around in four weeks once it had a dedicated programme manager working inside that structure. Not a framework doctrine. Governance and visibility.

 

The Real Adoption Curve

That pattern shows up in the adoption numbers too. Hybrid delivery has grown 57% since 2020, while purely predictive approaches have fallen 24% over the past three years. The pattern behind that shift is simple: pure Waterfall and pure Agile were both answering questions the actual work wasn’t asking, and organisations are admitting it with their adoption numbers rather than in a strategy memo.

Rigid, one-size-fits-all implementation is what actually ages badly, more than the framework choice underneath it. Portfolios that mix delivery models by what the work actually demands, a predictable cadence for mature products, flow-based delivery for continuous work, fast validation cycles for early bets, consistently outperform portfolios that force every initiative through the same certified process regardless of fit.

 

What This Means for the Programme You’re Running Now

The practical shift isn’t abandoning structure for chaos. It’s building governance that travels with whichever delivery model fits the work, instead of assuming the delivery model is the governance.

Start by naming decision rights before the kickoff, not after the first dispute breaks something. Someone owns scope changes. Someone owns the call when a dependency slips. Everyone on the programme should be able to name both without asking. Build the escalation path into the plan itself rather than inventing one under pressure in week eight, and match the delivery model to the type of work in front of you rather than to what worked on the last programme. A regulated, multi-vendor transformation and a ten-person product team shipping a new feature are not the same problem, and forcing them through the same framework produces the same failure pattern twice.

Track what the framework was supposed to deliver, not whether the ceremonies happened. A team that ran every stand-up and still shipped nothing useful followed the process and failed anyway.

 

The Question Every Kickoff Should Answer First

Before the next programme gets a charter and a framework stamped on the cover page, ask a harder question first: who owns the decision when priorities collide, and does that answer exist in writing before the first sprint starts?

If it doesn’t, the framework on the cover page was never going to save the programme underneath it.

The Governance Training Nobody Budgets For

Nobody has ever failed a project because they didn’t know how to build a Gantt chart. However some failed because nobody taught them who was allowed to say no.

I have sat in quite a number of programme inductions, and every one of them covers the same ground. Scheduling. Budgeting. Risk logs. RAID templates. Reporting cadence. Then, somewhere around the second afternoon, someone puts up a slide about governance and the room’s attention visibly leaves the building. It is treated as the compliance module, the thing you sit through before you get to the real work. Nobody walks out able to say, with any confidence, who actually owns the call when a decision does not fit neatly on a template.

That gap is not an oversight. It is a choice organisations keep making, year after year, without ever naming it as one.

 

The curriculum has a hole in it

Ask a newly promoted programme manager to explain their RAID log and they will do it fluently. Ask them who has the authority to accept a risk above a certain threshold without escalating it, and watch the pause. That pause is the sound of someone realising they were never actually taught the answer, only ever expected to absorb it by watching more senior people long enough.

A 2026 AI Governance Gap Report surveying more than 500 HR professionals found that only 45 per cent of organisations provide AI literacy training to all employees, one concrete, measurable instance of the broader governance-literacy gap this piece is about. Two-thirds of HR teams are already using AI to shape compliance and policy decisions, and the same report found fewer than half have given employees even that AI-specific literacy training to keep pace. The gap shows up well beyond HR too, in every function that has ever built a decision-rights framework, filed it in a folder, and assumed the document itself did the teaching.

Documents do not teach. People do, and usually only the ones who were already senior enough to have picked it up somewhere else.

 

Decision rights are treated like folklore

Most organisations do not lack a governance framework. They have one, usually a good one, sitting in a policy library that almost nobody outside the PMO has opened. What they lack is a mechanism for turning that document into instinct.

The result is a workforce that learns decision rights the hard way: by guessing wrong in front of a steering committee, by escalating something trivial and being quietly told off for wasting everyone’s time, or by not escalating something serious and finding out only when it has become a crisis. Every one of those is an expensive way to teach a lesson that could have been taught in an afternoon.

This is where the “knowing-doing” research cited by Harvard Business Review becomes uncomfortable reading for anyone who runs a training budget. Two out of three managers say they are still uncomfortable having accountability conversations with their own people, despite most of them having sat through the leadership training designed to prepare them for exactly that. The problem was never a shortage of content. Knowing a framework exists and being able to act on it under pressure are two entirely different skills, and organisations keep training the first while assuming it produces the second.

Governance training suffers from the same fault line. Knowing there is an escalation policy is not the same as recognising, in the middle of a stressful Tuesday, that the decision in front of you is the one the policy was written for.

 

The training everyone skips because it looks obvious

There is a reason this particular gap survives budget reviews when almost nothing else does. Governance training looks like it should be simple, so nobody prioritises building it properly. Everyone assumes the framework document is self-explanatory, right up until the moment someone makes the wrong call and the post-incident review discovers that three different people had three different understandings of who was supposed to decide.

I have run those reviews. The finding is almost never “the framework was wrong.” Almost always, nobody had ever been walked through what the framework meant in a live situation, so everyone applied their own version of common sense, and common sense is not actually common.

 

What actually needs teaching

A longer policy document will not fix this. Longer documents get read less, not more. The fix is teaching people to recognise a decision point before it arrives, not after.

That means running people through real scenarios instead of abstract categories, trading “what is your escalation threshold” for “here is a supplier problem that looks small and is not, what do you do in the next ten minutes.” It means naming, out loud and often, the handful of decisions in your organisation that carry disproportionate weight, so people learn to feel the shape of one before it is labelled for them. And it means treating governance literacy the way you would treat safety training: refreshed, tested, and taken seriously enough that senior leaders visibly participate in it themselves, rather than left as a one-off induction module.

The organisations that get this right build fewer, better decision-makers rather than thicker governance frameworks: people at every level who can spot a genuine decision point on instinct, the same way an experienced engineer can hear an engine fault before the dashboard lights up.

You cannot budget for the crisis a decision creates and then refuse to budget for teaching people to see it coming.

 

 

Pre-Mortem: The Liability Chain Medicare’s AI Prior Auth Model Has Not Drawn

On 1 January 2026, the Centers for Medicare and Medicaid Services in USA launched the WISeR model in six states, introducing prior authorisation to procedures that traditional Medicare had always provided without it. Contracted companies now assess medical necessity using AI. Human clinicians are required to sign off on any denial. The Senate voted 46-50 in July 2026 to keep the programme running. One question has not been answered.

This is the twelfth piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

CMS is wagering that AI-assisted prior authorisation reduces unnecessary Medicare spend without producing the patient-safety incident that forces a political reversal. If WISeR delivers measurable waste reduction without a documented causal chain from AI denial to patient harm, it becomes the template for prior authorisation across Medicare nationally. If it produces that chain, a documented line from AI recommendation to denial to patient harm, it does not just end WISeR. It becomes the reference point that makes AI prior auth politically untouchable in federal health programmes for a generation.

 

The Assumption

CMS has answered every operational question about WISeR except this one:

When an AI recommendation leads a contracted clinician to deny care and a patient is harmed as a result, where does liability sit?

The model design places a human clinician between the AI output and the denial decision. That establishes a paper trail. It does not establish a liability framework. Contractors earn between 10 and 20 per cent of the savings generated by denials and lose that payment when a denial is overturned on appeal. That is a commercial penalty, not a clinical one. The Federal Tort Claims Act does not cover contracted entities. No federal court has tested whether a contracted clinician reviewing AI recommendations at volume carries the same duty of care as a treating physician making an independent clinical judgement.

The assumption doing all the work in this model is that the human review layer is accountability enough. That assumption has not been tested.

 

The Sequence

1 July 2025. CMS published the WISeR notice in the Federal Register and did not submit it to Congress under the Congressional Review Act. That omission would matter later.

1 January 2026. WISeR launched in New Jersey, Ohio, Oklahoma, Texas, Arizona, and Washington.

17 March 2026. The Washington Post published an exclusive: Medicare’s new AI gatekeeper was delaying care for seniors. The University of Washington’s medical system had nearly 100 patients waiting for epidural injections. In Arizona, Phoenix pain specialist Dr Matthew Crooks told Medscape that every epidural injection submitted in the first three months had been denied and described the system as completely nonfunctional and unsustainable. In Texas, initial AI approval rates ran at 62 per cent, against a 92 per cent national approval rate across Medicare Advantage.

25 March 2026. The Electronic Frontier Foundation filed a FOIA lawsuit against CMS in federal court in California, seeking records on WISeR’s AI algorithms, training data, bias safeguards, and the financial incentives paid to contractors. The suit confirmed that CMS had not made its AI methodology or vendor compensation structure publicly available seven weeks after launch.

6 April 2026. CMS published a Federal Register notice delaying prior authorisation implementation for certain services within the model to allow additional time for operational readiness. CMS also issued a corrective action order against one of its AI contractors. Both confirmed that the model’s operational design had not performed as intended in the first quarter.

12 May 2026. The Government Accountability Office issued its determination: WISeR met the Administrative Procedure Act definition of a rule and was subject to the Congressional Review Act. CMS had not made the required submission to Congress before the model took effect.

20 May 2026. Senator Ron Wyden and Representatives Suzan DelBene and Greg Landsman introduced resolutions of disapproval in both chambers, seeking to repeal WISeR under the CRA.

6 July 2026. Gold carding launched in Washington state. Providers achieving a 90 per cent affirmation rate across a minimum of ten prior authorisation requests become exempt from further review for covered services. Quarterly rollout to the remaining five states is planned.

16 July 2026. The Senate voted 46-50 against advancing the disapproval resolution. Party line. WISeR survived. The liability question the GAO had exposed survived with it.

The Pager

Dr Mehmet Oz, Administrator of the Centers for Medicare and Medicaid Services.

The message: WISeR’s accountability chain has not been drawn. The model places a contracted clinician between an AI denial recommendation and a Medicare beneficiary, but no published document establishes where negligence sits when a patient is harmed following an AI-assisted denial. The Federal Tort Claims Act does not cover contractors. Contractors point to the human clinician. Clinicians are reviewing AI output under volume pressure with no published duty-of-care standard for that specific context. When the first federal lawsuit tests this configuration, and one will, CMS will need a published framework, not a contract clause. That framework is easier to write before litigation than after.

 

The Proof

Gold carding is the model’s self-correction mechanism. If quarterly rollout reaches all six states and the 90 per cent affirmation threshold functions as a genuine quality signal, the AI layer contracts over time as trust is established. Proven providers exit prior auth. New entrants face the review. The model becomes calibrated rather than blanket.

If gold carding stalls or rollout criteria are applied inconsistently across jurisdictions, the AI layer expands without a release valve. Prior auth burden accumulates regardless of provider track record. The model becomes a cost-reduction instrument with no exit for providers who have earned one.

The proof of the bet is not the aggregate savings figure. It is whether WISeR, by the end of 2026, has published a liability framework and delivered gold carding in all six states. Without both, the model is running on the same untested assumption it started with.

 

Verdict

If CMS publishes a liability framework for AI-assisted denials before a federal case forces the question, and gold carding delivers consistent rollout across all six states, WISeR will be the strongest government evidence yet that AI-assisted utilisation review can reduce Medicare waste without a patient-safety crisis. The accountability design would become the reference for every federal health programme that follows.

Without the liability framework, WISeR accumulates its risk quietly. Not through a single dramatic incident, but through the gap between AI recommendation volume and human review capacity, compounded by an accountability vacuum no published document has yet closed. That gap does not stay open indefinitely.

Pre-Mortem: The US Government’s 3,611 AI Use Cases

On 3 April 2025, the White House issued OMB Memorandum M-25-21, directing every major federal agency to appoint a Chief AI Officer, expand the use of artificial intelligence across government operations, and manage risk proportional to each system’s impact on citizens. Twelve months later, the Federal Agency AI Use Case Inventory records 3,611 AI use cases across 56 agencies, more than double the prior year’s total. A May 2026 survey of more than 200 technology executives across civilian and defence agencies found 53% are actively planning agentic AI pilots. Only 8% of those agencies have incident response frameworks in place.

This is the eleventh piece in the Pre-Mortem series. Five questions, applied to the public record, before a programme has had the chance to succeed or fail.

 

The Bet

The US government is betting that embedding AI across 3,611 federal workflows covering benefits decisions, immigration adjudications, healthcare determinations, and law enforcement, will make government faster and more efficient before the accountability architecture governing those decisions is clarified. OMB M-25-21 requires Chief AI Officers, public AI inventories, and risk management proportional to impact. The hard compliance deadline for role-specific AI training arrives in September 2026. If that architecture catches up to the deployment before a consequential wrong decision reaches a citizen with no named relief, the bet holds.

 

The Assumption

The expansion’s credibility turns on one unanswered question: whether the Federal Tort Claims Act, designed to govern negligent acts by human federal employees, applies without amendment to decisions made by AI agents running inside federal systems. The same May 2026 survey found only 44% of agencies include vendor liability clauses in AI contracts, and only 29% have documented kill-switch procedures. The legal architecture governing accountability in federal government was designed for humans acting on behalf of the state. No court has ruled on whether it extends to the agents they built.

 

The Sequence

In 2024, federal agencies reported 1,757 AI use cases. By 2025, that figure had grown to 3,611. In March 2026, the Department of Veterans Affairs expanded AI use in claims processing, with 215 of its 367 AI systems classified as high-impact, covering benefit eligibility, healthcare access, and fraud detection. In May 2026, the majority of agencies were planning agentic pilots, with only 20% having defined pre-deployment testing policies. The AI reached citizens before the accountability reached the AI.

 

The Pager

Russell Vought, Director of the Office of Management and Budget, carries the M-25-21 mandate at the centre of the federal AI expansion. Every covered agency has designated a Chief AI Officer responsible for inventory, risk management, and AI governance at agency level. The VA alone runs 215 high-impact AI systems. No single published document names what relief is available to a veteran whose claim was influenced by one of those systems, which official carries accountability for that decision, or whether the Federal Tort Claims Act applies when the acting party is software, not a civil servant.

 

The Proof

The measure that would settle this is a published legal standard: a named accountability chain clarifying who carries liability when a federal AI agent makes a consequential wrong decision, whether government, vendor, or joint, and whether the Federal Tort Claims Act applies or new legislation is required. No such standard has been published. OMB M-25-21 requires risk management proportional to impact. It does not name the relief available to a citizen when that risk management fails, nor the date by which that question must be answered.

 

Verdict

If OMB publishes, before the September 2026 training compliance deadline, a named accountability standard for AI-driven decisions in high-impact federal systems, covering who carries liability when the AI is wrong and what legal remedy a citizen holds, the expansion will stand as the most deliberate attempt the US federal government has made to govern AI before it reaches citizens at scale. Without that, the US government has put AI into 3,611 workflows and left the question of who carries the call when the AI gets it wrong to be answered in court, by accident, or not at all.