One Token, Three Tools Deep, 2,500 Companies Across Five Continents

One leaked automation token. Three tools deep. Over 2,500 organisations across five continents exposed. That is the entire anatomy of 2026’s largest AI supply chain breach, and it took roughly forty minutes to happen.

How a Security Scanner Became the Weapon

In March 2026, a threat group tracked as TeamPCP compromised Trivy, a widely used open-source security scanner, through a leaked automation token. From there the attackers force-pushed malicious code into Trivy’s own version tags. When LiteLLM, a popular AI gateway that routes traffic between applications and model providers, pulled the poisoned scanner into its build pipeline, it published two compromised releases to PyPI, versions 1.82.7 and 1.82.8. A hidden file inside those packages executed automatically the moment Python started, no import required, harvesting cloud keys, repository tokens, SSH credentials, Kubernetes secrets and AI provider API keys from anyone who installed it. The malicious packages were live for roughly forty minutes before removal. That was enough.

 

A Genuinely Global Exposure List

CloudSEK’s analysis, later corroborated by an FBI advisory in July, identified more than 2,500 affected organisations and 434,000 exposed CI/CD pipelines. The named companies span the world rather than one region: AWS, Nvidia, Salesforce and Airbus’s US Space and Defense division from the US, Samsung Electronics and MediaTek from Asia, Siemens, Volkswagen and Munich Re from Germany, Thales and Orange from France, Roche from Switzerland, Philips from the Netherlands, Vodafone and the London Stock Exchange Group from the UK, and Thomson Reuters from Canada, alongside Cisco, FedEx, Deloitte and dozens more. This was never a story about one country’s technology sector. It was a demonstration of how deeply a single open-source dependency now sits underneath enterprise infrastructure everywhere.

 

The Same Argument, a Different Layer

This publication has already made the case that The September AI Outage Had Two Real Causes. Almost Everyone Reported One. LiteLLM extends that argument one layer deeper, into the software supply chain feeding the infrastructure itself. A vendor risk register that lists which cloud provider or model vendor an organisation depends on is no longer sufficient. It also needs to account for the open-source scanners, build tools and CI pipelines every one of those vendors quietly depends on, tools most procurement processes never ask about because nobody signed a contract for them.

 

What a PMO Should Actually Check

Three questions belong on every technology steering committee agenda this quarter. First, does anyone in the organisation maintain a current list of the open-source build and security tooling embedded in critical CI/CD pipelines, not just the paid vendors. Second, when was a leaked or rotated credential last tested as an actual incident scenario, rather than a line item in a policy document. Third, if a dependency three layers removed from a primary vendor were compromised tomorrow, how long would it take anyone to notice. The organisations named in the LiteLLM breach were not careless. They were exposed by a dependency most of them did not know they had, which is precisely the risk a vendor register built only around visible, contracted suppliers will always miss.

The September AI Outage Had Two Real Causes. Almost Everyone Reported One.

On the morning of 3 September 2026, ChatGPT, Claude and Grok all failed within roughly ninety minutes of each other. Within hours, most coverage had settled on a single explanation: a shared Microsoft Azure failure had taken down three competing AI platforms at once. The postmortems that actually exist tell a different story, and the difference matters more than the outage itself.

 

What Each Company Actually Said

SpaceXAI was the most direct. The company confirmed a failure at its Memphis compute facility, apologising to what it called its “impacted compute partners,” language that implies other organisations lease capacity in the same facility. Grok was down for roughly three and a half hours. OpenAI gave its own, separate explanation: an engineer identifying as the incident commander stated on Hacker News that the cause was a routing error inside OpenAI’s own infrastructure, surfacing roughly ninety minutes after SpaceXAI’s Memphis failure and unconnected to the GPT-6 Astra launch that followed the next day. Anthropic never published a technical cause at all. It confirmed elevated errors across several Claude models starting around 9:26am ET and reported the issue resolved, but the actual mechanism behind it was never disclosed.

 

Why the Azure Theory Doesn’t Hold Up

The Azure theory spread fast because it was the simplest explanation available in the first few hours, and it fit an existing assumption that most enterprise AI risk registers already carry: if these vendors compete, their infrastructure probably doesn’t overlap. It did not hold up. Microsoft logged no Azure incident for that window, and Cloudflare said in a statement that it was not experiencing any significant service disruptions. The timeline argues against one shared trigger too, SpaceXAI’s and Anthropic’s incidents began four minutes apart from each other, and OpenAI’s followed roughly ninety minutes later, a pattern that fits three separate incidents more comfortably than one shared one.

 

The Question Nobody Answered

The part worth noticing is not that the Azure story was wrong. It is that the two companies which did give an explanation gave two different ones, and the one that stayed quiet is the one whose incident opened four minutes before SpaceXAI’s, the tightest timing overlap of the three. Nobody, including Anthropic, ever confirmed or ruled out whether Claude’s infrastructure touches the same Memphis facility SpaceXAI apologised to its partners about. A vendor risk register cannot verify a question its own vendor never answers.

None of this means multi-vendor AI strategy is worthless. It means the diversification most risk registers assume is rarely something a customer can actually verify from the outside, and the one moment it gets tested, a real outage, is also the moment vendors are least likely to disclose the detail that would let you check. Two of three companies gave a specific, different cause. One gave none. Whichever of those a programme is depending on, the honest entry in a vendor risk register reads “unconfirmed,” not “diversified.”

SAP Told Its Customers Which AI to Use. Three Per Cent Listened.

By February 2026, DSAG’s own Investment Report had already surfaced the gap SAP would spend the rest of the year trying to close by other means. Only 3 per cent of surveyed companies ran their production AI use cases on SAP’s own tools, while 77 per cent ran theirs on non-SAP solutions instead, Microsoft Copilot chief among them. Two months later, SAP rewrote its API policy. Section 2.2.2 of the new terms blocks external AI agents from calling SAP systems directly, prohibiting what it calls independent scheduling or execution of API calls. Every agentic use case now has to route through Joule, SAP’s own assistant, adding a second layer of inference and cost to anything a customer’s AI tools want to do inside SAP.

 

Two Competitors Bet the Opposite Way

Salesforce and ServiceNow read the same gap differently. Salesforce launched Headless 360 in April 2026, giving external agents direct access through REST, Model Context Protocol tools and the command line, no mandatory detour through a proprietary assistant required. ServiceNow’s Action Fabric, announced at Knowledge 2026, opened two decades of workflows and business rules to any agent through the same open standards. Neither company built a legal wall around its intelligence layer. Both are betting that customers will choose their AI because it is good, not because the contract requires it.

 

The Mandate Had a Structural Problem Too

Joule is only available to customers on RISE or GROW with SAP cloud contracts, which shuts out the thousands of organisations still running on-premise ECC. Support for ECC 6.0 ends in 2027, so over ten thousand customers are being asked to migrate their core system and adopt a mandated AI layer on the same timeline, set by the vendor rather than the customer. SAP’s own chief customer officer has disputed the DSAG survey’s reach, noting that fewer than 100 organisations responded, which is a fair challenge to the precision of the number. It does not explain why SAP is separately funding a hundred-million-euro partner investment programme to accelerate adoption of a product it says the market already wants.

 

The Hedge That Followed the Mandate

At Sapphire 2026, a few months after the API policy took effect, SAP quietly hedged its own bet. Alongside Joule Studio 2.0, it introduced an AI Agent Hub, described as vendor-agnostic governance for SAP and third-party agents alike, bundled into the platform at no extra charge. A company confident that mandating Joule had closed the gap DSAG measured in February would not need to build the open door it spent the spring telling customers they did not need.

 

The Question Every PMO Should Be Asking Now

None of this is really about SAP. Every PMO running a platform strategy right now is making the same call in miniature, whether to route AI agents through one sanctioned tool or leave the door open for whichever tool the business actually adopts. SAP had the market power to write the mandate into a contract, and it wrote that mandate only after its own user group had already published the adoption numbers it was meant to fix. Assume your organisation has less leverage than SAP did, and plan the exit clause before you write the mandate, not after the adoption numbers come back.

Your AI Steering Committee Meets Monthly. Your AI Ships Weekly.

Most AI governance committees meet on a monthly or quarterly cadence. Most AI systems do not wait for them.

Anthropic’s own release notes tell the story in a single scroll. Opus 5 launched in late July 2026, a step-change release over Opus 4.8, which was barely two months old at the time. Sonnet 5 landed a month before that, at the end of June. Opus 4.8 before that, in late May. Opus 4.7 before that, in April, with breaking API changes attached. Opus 4.5 in November 2025. Sonnet 4.5 in September. Six major model versions inside ten months, with incremental API and feature changes logged in between on a near-weekly basis. That is one vendor’s changelog. Most enterprises are running several.

 

The Governance Side Has Not Kept Pace

Deloitte’s 2026 State of AI in the Enterprise survey, covering 3,235 business and IT leaders across 24 countries, found that only 21% of organisations have a mature governance model in place for agentic AI. Meanwhile, 74% of the same leaders expect their organisation to be using AI agents at least moderately by 2027, with 23% expecting extensive use and 5% expecting full integration.

That is not a small gap closing gradually. It is deployment intent running well ahead of the governance capable of managing it, at the exact moment agentic systems are gaining the authority to act rather than just draft.

 

People Have Already Stopped Waiting

The gap does not sit quietly while committees catch up. A Cybernews survey of more than 1,000 US employees found that 59% use AI tools their employer never approved. Among executives and senior managers specifically, the same survey found 93% do the same. The people who would sit on the steering committee approving AI use are, by a wide margin, the people going around it.

That is the quiet part of this problem. The people with the authority to slow things down are the ones setting the pace of the tools rather than the pace of the committee.

 

What Happens When the Gap Surfaces

Gartner’s prediction for next year is specific: 40% of enterprises will have their autonomous AI efforts partly derailed by governance gaps discovered only after a production incident, not before one. Sanchit Vir Gogia of Greyhound Research, commenting on the same data, put the mechanism plainly: “the real governance problem is not model intelligence. It is delegated operational authority moving across trust boundaries faster than enterprises can instrument, constrain, or audit it.” His own warning, in his words: “Do not scale agents faster than you can govern their authority. A small number of well-governed agents will create more enterprise value than a sprawling estate of clever, fragile, over-permissioned digital apprentices.”

 

What Governance Cadence Actually Needs to Look Like

The fix is not more meetings. A committee that reviews AI deployment quarterly cannot govern a system that changes weekly, no matter how thorough each quarterly review is. What changes the equation is moving from periodic review to continuous, tiered oversight: automatic flags for any new agent or capability change above an agreed risk threshold, standing authority delegated to a smaller operational group empowered to act between full committee meetings, and a real audit trail that lets the full committee review what was approved at pace, after the fact, rather than approving everything before the fact and creating the exact bottleneck this problem describes.

Keeping oversight in place does not mean keeping the quarterly cadence that was designed for a technology that no longer exists.

 

The Question Worth Asking Before the Next Steering Committee Meeting

Ask what actually changed in your AI environment since the last meeting, specifically, not generally. If nobody in the room can answer that with real detail, the governance structure is reviewing a snapshot that was already out of date the moment the meeting started.

The organisations that get this right are not the ones meeting more often. They are the ones who redesigned what needs a meeting at all, and gave someone standing authority to handle everything else in between.

EU Cyber Resilience Act – Reporting Duty Begins on an Unfinished Platform

On 11 September 2026, the European Union’s Cyber Resilience Act began requiring manufacturers of products with digital elements to report actively exploited vulnerabilities and severe incidents within 24 hours of becoming aware of them. The duty runs through the EU Agency for Cybersecurity’s new Single Reporting Platform, which ENISA switched on the same day while describing it as having reached only “initial operating capability.” That the deadline and its supporting infrastructure arrived on the identical date is itself the story, and it comes with a genuine credit: a manufacturer selling into the EU can now file one exploited-vulnerability report and have it reach every relevant national authority, rather than notifying each member state separately.

This Pre-Mortem asks the questions a post-mortem would ask, before failure is possible: what is being bet on, what single assumption could break it, what got decided before the safeguards existed, who carries the pager when it fails, and what proof would settle whether it worked. It is the diligence a compressed rollout deserves before its first real test, not after.

 

The Bet

The EU is betting that centralising exploited-vulnerability and severe-incident reporting through one platform gives ENISA and national CSIRTs (Computer Security Incident Response Teams) EU-wide visibility into what is actually being attacked, and that manufacturers will treat a 24-hour clock as workable even while the tool underneath it is still being built out. The bet favours momentum over completeness. The European Commission’s own guidance on the CRA’s reporting obligations confirms the platform was already operational on 11 September 2026, the same day the 24-hour duty took effect, rather than waiting until it was finished. It also asks industry to build compliance discipline around a tool still being assembled, in the same quarter that discipline is tested.

 

The Assumption

The load-bearing belief is that a “single” platform stays meaningfully single even with core pieces missing. As the specialist tracker cyberresilienceact.eu reported on launch day itself, voluntary reporting under Article 15, a programming interface, and a field recording exactly when a manufacturer became aware of an incident all “did not arrive with it,” and access runs only through an Assigned Representative role via EU Login multi-factor authentication. If that gap persists, manufacturers with products across many jurisdictions and limited access routes could end up filing manually into individual CSIRTs anyway, the exact fragmentation Article 16 was written to end.

 

The Sequence

The Commission did not settle what counts as a manufacturer becoming “aware” of a reportable event, the trigger that starts every clock in the regime, until its guidance published on 27 July 2026, six weeks before the duty began. For most of the preceding compliance-planning year, manufacturers had no published definition of the one moment that decides whether a 24-hour window has even opened. ENISA’s own Assigned Representative registration guidance still carried dated updates in launch week itself, on 9 and 10 September 2026, the guidance for actually using the platform arriving alongside it rather than well ahead of it.

 

The Pager

ENISA’s Executive Director, Juhan Lepassaar, put his name to the launch, framing it as a step toward “a more resilient Digital Single Market,” and carries the operational pager for the platform itself. Above him sits European Commission Executive Vice-President for Tech Sovereignty, Security and Democracy Henna Virkkunen, who holds the political brief for EU cybersecurity policy, having worked on the Cyber Resilience Act dossier in Parliament. Beneath both, national market surveillance authorities enforce the penalty tier attached to Article 14 failures: up to €15 million, or 2.5 per cent of global turnover. No individual has been named as accountable for the missing programming interface and voluntary-reporting channel, only an agency-level pledge to keep improving.

 

The Proof

No public dashboard, uptime figure, or notification-volume metric has been attached to the platform’s launch, and ENISA’s own pledge to keep improving it carries no completion date. The measure that would actually settle the question, whether a report filed with one coordinating CSIRT reaches every other relevant CSIRT and ENISA at the speed the platform promises, has not been made public in any testable form. That will only become visible once a real, cross-border, multi-product incident runs through the system at volume, rather than from the demonstration traffic of launch week, and eighteen months allows roughly two CRA reporting cycles for that test to arrive.

 

Verdict

If ENISA closes the Single Reporting Platform’s programming-interface and voluntary-reporting gaps, and attaches a public date to doing so, before the platform faces its first genuinely cross-border, multi-CSIRT incident, then 11 September 2026 will read as a pragmatic phased launch that got the hardest deadline live on time. If those gaps are still open when that incident arrives, manufacturers already working to a 24-hour clock will discover that the single reporting platform was, in practice, still several platforms wearing one name.

A Tenfold Attack Surge in London, a Governance Mandate in Abu Dhabi

In January to May 2026, UK healthcare providers logged 264,000 attack events, against 27,000 for the whole of 2025. That is a tenfold jump in five months, according to SonicWall’s telemetry across NHS-linked sensors, and it occurred on a sector where the gap between what boards think they know and what is actually protected has never been wider.

 

The Long Tail of a Single Attack

The clearest illustration of what that gap costs sits two years in the past and is still unresolved. The June 2024 ransomware attack on Synnovis, the pathology provider serving South East London hospitals, exposed data from roughly a million NHS patients and forced more than 10,000 outpatient appointments and 1,700 elective procedures to be postponed. As of April 2026, South London and Maudsley NHS Foundation Trust was still processing pathology results without fully restored systems, and the incident has been linked to a patient death at King’s College Hospital. Nearly two years is not a recovery timeline any board signed off on. It is what happens when governance treats cyber resilience as an IT programme rather than a clinical safety issue.

 

Reactive Versus Mandated: Two Governance Models

The UK’s surge sits inside a largely reactive governance model, where boards respond to incidents and regulators tighten guidance after the fact. Abu Dhabi has taken the opposite route. Its Healthcare Information and Cyber Security standard, now in its second version, is built around six pillars, and governance is listed first, ahead of resilience, capability, partnerships, maturity and innovation. Every hospital, insurer and medical device maker operating in the emirate has to comply, and the standard explicitly frames cyber security as an organisation-wide responsibility covering people and process rather than a narrow technology control bolted onto IT. It is a mandated structure built before the incident, rather than a review commissioned after one.

 

Why the Stakes Keep Rising

The economics make the governance question harder to defer. Healthcare ransom demands have climbed sharply, with one widely cited industry analysis putting the average demand at $16.9 million, up from $577,800 the prior quarter, and healthcare remains one of the most targeted sectors globally, with 77 per cent of organisations reporting a ransomware attempt in the past twelve months, because attackers know disrupted care creates leverage no other industry carries. A board that treats a green status on a cyber dashboard as sufficient assurance is trusting a number it has usually never interrogated.

 

What Boards Should Actually Be Asking

Two different regulatory paths, one shared conclusion. Cyber governance in healthcare cannot sit exclusively with a CISO or an IT director reporting up through a technical channel. It needs a named board-level owner who can answer, in plain terms, three questions: which clinical services stop if this system fails, how long the organisation can run on manual process before patient safety is affected, and when that assumption was last tested rather than assumed. The Synnovis timeline suggests most boards do not currently have confident answers. The Abu Dhabi model suggests what building toward one actually looks like, governance treated as the first pillar of resilience rather than the last line of a post-incident report.

Pre-Mortem: The UN’s Global Dialogue on AI Governance

On 6 July 2026, the United Nations opened its first General Assembly-mandated forum on AI governance in Geneva. All 193 member states attended, the first time every country, developing and developed alike, has held a formal seat at an AI governance table. Secretary-General António Guterres named four priorities: common safety standards, human-rights red lines, capacity-building for developing nations, and environmental transparency. After two days, the forum closed with a co-chair summary. Not a treaty. Not an enforcement mechanism. The document records what governments agreed in the room. It does not bind any of them to act on it.

A pre-mortem applies five fixed questions to a public commitment before the outcome is known: what is being bet on, what single assumption underpins it, what was decided before governance existed, who carries it if it fails, and what would prove it worked. This piece applies that format to the public record of the first UN-mandated global AI governance dialogue, convened in Geneva on 6 and 7 July 2026.

 

The Bet

The UN’s bet is that convening all 193 member states repeatedly, Geneva first, New York in May 2027, produces a governance architecture capable of managing AI at global scale before the harms it is designed to prevent have already arrived. The mechanism is norm-setting: shared principles, common language, multilateral dialogue, with no enforcement powers and no treaty obligations. The independent scientific panel, co-chaired by Yoshua Bengio and Maria Ressa, published its first report on 1 July 2026, before the forum opened. Guterres proposed a Global Fund for AI and an AI Child Safety Pledge. Those are specific commitments, publicly made. The bet is that naming four priorities in a room of 193 governments changes what those governments do next.

 

The Assumption

The Dialogue’s credibility turns on a single calculation: that member states, with vastly different AI capabilities, legal systems, and strategic interests, will treat non-binding shared principles as a meaningful constraint on sovereign AI deployment decisions. The Dialogue produces a co-chair summary, not a resolution or treaty. A state that attends, agrees on human-rights red lines, and then deploys AI in ways that cross them faces no published consequence. At the Dialogue itself, the United States delegation argued that voluntary cooperation between government and industry, not binding rules, is the only approach agile enough to govern AI, a preference for exactly the model the Dialogue is testing. The architecture assumes that participation changes behaviour.

 

The Sequence

In 2017, Guterres called for global AI governance. In September 2024, the General Assembly adopted a Global Digital Compact that included AI provisions. In August 2025, the General Assembly established the Global Dialogue by resolution. The first session convened in July 2026. In the same period, the largest frontier AI models were trained, deployed at scale, and embedded in healthcare, finance, law enforcement, and defence, across countries that attended Geneva and agreed that safety is a priority. The Bengio-Ressa panel stated plainly that science currently cannot guarantee AI will not cause catastrophic harm as capabilities increase. The governance structure followed the deployment decisions. That is the sequence.

 

The Pager

Ambassador Egriselda López of El Salvador and Ambassador Rein Tammsaar of Estonia co-chaired the first session. Amandeep Singh Gill, UN Special Envoy for Digital and Emerging Technologies, coordinates the process. When a member state deploys an AI system that crosses the Dialogue’s own human-rights red lines, no document names the consequence. The co-chair summary records what was agreed. It is not a mechanism. The next session is May 2027 in New York. The question of who carries accountability between now and then has no published answer.

 

The Proof

Guterres named four priorities at Geneva. The measure that would prove any of them produced outcomes by May 2027 is at least one of these: a common safety standard that member states have adopted, a documented case where a human-rights red line prevented a harmful deployment, a committed capacity-building fund with named recipients, or an AI environmental reporting mechanism with verified data. The co-chair summary is the output the Dialogue committed to producing. An outcome requires a different commitment entirely.

 

Verdict

If the second session in New York, May 2027, produces a named accountability mechanism for at least one of Guterres’s four priorities, with a state or institution publicly committed to carrying it, the Geneva forum will stand as the first step of something with teeth. Without that, it stands as the moment 193 governments agreed that AI is the most consequential technology of the era, and scheduled a follow-up.

Your Reputation Travels Faster Than You Do. Act Accordingly.

Most executives manage their reputation like a local matter: how you’re seen in this room, on this team, in this market. That’s the wrong frame. Reputation moves through networks faster and further than any individual career move, and it arrives in the next room before you do.

 

The Research Behind Why Word Travels

A 2022 study in Science, based on five years of randomised experiments across 20 million LinkedIn users, 2 billion new connections, and 70 million job applications, found that professional information travels most efficiently through weak ties, not close friends. The loose, wide network of people who know you a little, rather than the small circle who know you well, is what actually carries information about you into rooms you haven’t entered yet.

That mechanism cuts both ways. It is exactly how good work gets you noticed somewhere new. It is also exactly how a reputation for cutting corners, mistreating people, or leaving a mess behind you gets there first.

 

What Happens to Reputation That Travels Badly

The clearest, most rigorously measured evidence of this comes not from executive search literature, which is surprisingly thin on hard numbers, but from corporate governance research on company directors. A 2005 study in the Journal of Accounting Research tracked 409 US firms that restated their earnings between 1997 and 2001. Directors of those firms lost roughly a quarter of their positions on other, unrelated companies’ boards afterward, not just the one where the restatement happened. A related 2007 study in the Journal of Financial Economics found that outside directors named in shareholder fraud lawsuits saw a measurable decline in how many other directorships they held, even at companies with no connection to the original case, at an estimated cost of roughly $1 million per lost seat.

That is reputation travelling, quantified: conduct in one boardroom measurably closing doors in boardrooms that had nothing to do with it.

 

Real Cases, Not Hypotheticals

Steve Wynn resigned from Wynn Resorts in 2018 following sexual misconduct allegations. The consequences did not stay in Nevada. Massachusetts gaming regulators, investigating a market he had never previously operated in, fined the company $35 million and forced it to strip his name from its brand-new $2.6 billion property before it opened, renaming Wynn Boston Harbor to Encore Boston Harbor specifically to distance the business from him. He personally paid $10 million in 2023 to permanently exit the Nevada gaming industry. A reputation formed in one state travelled into a state where he had never done business, and shaped how a market he’d never worked in treated him before he arrived.

Travis Kalanick’s departure from Uber followed him into an entirely new, unrelated venture years later. Coverage of his food-delivery startup CloudKitchens traces his Uber exit “amid a firestorm of privacy concerns, allegations of widespread sexual harassment and gender discrimination,” then quotes a former CloudKitchens executive calling it “the most toxic place I’ve ever seen or experienced,” and an operator who said the company “tried to destroy” the brand he had built there. The new business was never assessed purely on its own merits. It was read through the lens of the one he had just left.

Not every case ends the same way. Andreessen Horowitz invested $350 million in Adam Neumann’s new venture Flow in 2022, valuing it above $1 billion before it had launched, despite WeWork’s collapse from a $47 billion to an $8 billion valuation under his leadership. Marc Andreessen’s public justification leaned on second chances: “we love seeing repeat-founders build on past successes by growing from lessons learned.” Reputation still shaped every headline, every term, and every question asked about the deal, even though it never blocked the capital.

 

Acting Accordingly

One caveat is worth stating directly: nobody has produced a clean statistic for how much weight boards or recruiters place on informal, back-channel reputation versus formal references. That data mostly doesn’t exist, and anyone who claims otherwise is making it up. The mechanism, though, is well documented: wide, weak professional networks carry information fast, and reputational damage in one role measurably reduces opportunity in entirely unrelated ones.

The practical implication isn’t paranoia. It’s that the version of you that shows up in a room you’ve never been in was written by people you may not remember meeting, months or years before you walked in. Act like the story is already there, because it usually is.

Pre-Mortem: The Big Four’s AI Citation Problem

On 28 July 2026, PwC Middle East responded to an investigation into four of its own published reports. The investigation, run by the AI-detection company GPTZero, had found fabricated citations, non-existent sources, and, in one report, a teenage blogger with 280 followers cited as an authority on a JPMorgan initiative. PwC’s statement: the company “takes the accuracy of our published research seriously” and was “updating a limited number of supporting citations.”

PwC was not first. It was the fourth.

This is the sixteenth piece in the Pre-Mortem series. Five questions, applied to the public record, before the outcome is known.

 

The Bet

Deloitte, EY, KPMG and PwC are betting that a pattern spanning five publicly documented reports, four countries, and under two years can be absorbed as unconnected incidents rather than treated as a shared problem with a shared cause. Each firm has responded on its own terms: a partial refund from Deloitte, quiet withdrawals from EY and KPMG, a promise to update “a limited number” of citations from PwC. None has published a shared verification standard. None has described what changes in how AI-assisted work is reviewed before the next report carries its name. The bet is that four reputations, built over more than a century, can absorb five independently verified failures of the most basic check a research report is supposed to pass: that the sources it cites exist.

 

The Assumption

Every one of the four firms has offered a version of the same explanation once caught. KPMG cited guidelines requiring human oversight to validate content and verify sources. PwC cited quality control processes it expects all its people to adhere to. The assumption underneath both statements: that a written guideline is itself a control, that if a policy exists, a human somewhere is presumed to have applied it before publication. EY’s report, “Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems,” carried the names of two partners and a senior manager in its byline. GPTZero’s analysis put the document at roughly 72 per cent AI-generated content, with more than half its 27 sources failing to correspond to anything real. Two partners and a senior manager reviewed that document before it went out, in name. What “reviewed” required in practice is the question none of the four firms has answered.

 

The Sequence

December 2024. PwC Middle East publishes “Agentic AI: The New Frontier in GenAI,” later found by GPTZero to contain fabricated citations.

October 2025. KPMG publishes “Total Experience: Redefining Excellence in the Age of Agentic AI.” GPTZero later finds 45 citations, 5 accurate, at least 16 fabricated.

October 2025. Deloitte refunds AU$97,000 of its A$440,000 contract with Australia’s Department of Employment and Workplace Relations, after a fabricated Federal Court quote and references to non-existent research papers are identified.

November 2025. Newfoundland and Labrador’s C$1.6 million Deloitte health workforce report is found to contain fabricated citations, including one crediting a Dalhousie University researcher as author of a paper that does not exist. Premier Tony Wakeham calls it “concerning.” Deloitte stands by its findings.

27 April 2026. South Africa’s draft National AI Policy is withdrawn 17 days after publication, after 6 of 67 citations are found fabricated. Minister Solly Malatsi calls it “an unacceptable lapse.”

14 May 2026. EY withdraws “Points of Attack” after GPTZero finds more than half its 27 sources do not correspond to real material.

12 June 2026. GPTZero publishes its investigation into KPMG. Five days later, this series covers a separate KPMG story without connecting the two.

28 July 2026. GPTZero publishes its investigation into four PwC Middle East reports. PwC responds that it is updating “a limited number of supporting citations.”

 

 

The Pager

Five public failures, four countries. Three were identified by the same three researchers, Paul Esau, Om Ogale and Alex Cui, working at GPTZero, not at any of the firms and not at any client who paid for the work. Every firm-level response has stopped at the firm: a refund, a report removed from a website, a statement that guidelines exist. No named individual at any firm has been identified as responsible for approving a document whose sources were not real. The one structural change on record did not come from a firm. Newfoundland and Labrador overhauled its own procurement process, requiring disclosure of AI use in future contracts. The government fixed what the contractor did not.

 

The Proof

None of the four firms has published a verification standard: a description of what checking a citation actually involves before a report carries its name. That is the proof measure, not an apology and not a quiet correction, but a public description of the review step, specific enough to be checked against the next report. The IAASB, the International Auditing and Assurance Standards Board, is revising ISA 500, the international standard governing what constitutes sufficient, appropriate audit evidence. That project is still at the research stage and covers formal audit engagements, not the thought-leadership publishing where three of these five failures occurred. Until one firm publishes what verification looks like in practice, every new report each of them publishes resets the same test.

 

Verdict

If one firm publishes a specific, checkable verification standard before a sixth incident surfaces, it becomes the reference point the other three are measured against, the position peer accountability once created around data breach disclosure, where one actor’s transparency made silence from the others harder to sustain. Newfoundland’s government has already shown the structural fix is available: a procurement clause requiring AI disclosure, written in days. If no firm moves first and a sixth incident surfaces, the pattern stops reading as isolated mistakes and starts reading as an industry’s operating baseline. Five failures in under two years, three caught by the same outside team. The firms selling AI governance advisory to clients have not yet demonstrated they can apply the same standard to their own published work. The next report each of them publishes is the test.

EasyJet Fixed an Age Bias in Recruitment. Most Digital Transformation Teams Haven’t

 

The number of easyJet cabin crew aged over 50 has more than doubled since 2022, up 127 per cent, according to the airline’s own figures. EasyJet says crew aged over 60 have “almost quadrupled” over the same period, and the airline has opened a fresh recruitment drive for the 2027 flying season, with applications opening in September. Getting there took a deliberate campaign. A large share of potential applicants assumed cabin crew work was reserved for younger people, and easyJet’s Director of Cabin Services, Michael Brown, put the fix plainly: over-50s bring both the skills to do the job and “a wealth of life experience that is appreciated by our customers and colleagues alike.”

 

This Is a Bigger Problem Than One Airline

The Centre for Ageing Better’s State of Ageing 2025 report shows why that perception carries a cost well beyond one airline. The UK’s employment rate for 55 to 64 year olds sits at 65 per cent, against 75 per cent in the Netherlands and Switzerland and 81 per cent in Iceland. The wider 50 to 64 employment rate sits 14 percentage points below the 25 to 49 rate. Closing that gap by 2030 would add an estimated £9 billion a year to the UK economy and £1.6 billion in annual tax and National Insurance revenue, according to the same research. That is the scale of value sitting behind a single, correctable assumption about who is fit to do a job.

 

The Same Bias, Earlier and More Expensive

The same assumption shows up earlier, and more expensively, in technology. CWJobs, working with the Centre for Ageing Better, surveyed 2,000 UK workers plus 250 people in tech who had experienced age discrimination, and found that tech employees start experiencing age bias at 29 and are considered “too old” by 38, roughly a decade before most people reach senior delivery roles. Forty-one per cent of tech workers report observing ageism at work, against 27 per cent across other sectors. Forty-seven per cent say they weren’t offered a role because of their age, and 31 per cent say they were passed over for promotion for the same reason. “Digital skills shortages mean discriminatory attitudes against age makes no business sense,” CWJobs director Dominic Harvey said when the findings were published, a point that has only got truer as the skills shortage he was describing has continued.

 

The Bias Has Already Reached a Tribunal

In Selazar Limited v McCabe, a tech company’s 29-year-old founder was found to have instructed a recruitment consultant to find “a younger team member who was more in tune with a young tech start company” in place of the firm’s 55-year-old finance director. The tribunal awarded her £125,604.98, including £20,000 for injury to feelings, and heard evidence that the founder had also signalled to potential investors that she was “too old to understand” the business. The case puts a figure on the same instinct easyJet had to overcome in reverse: treating experience as a cultural mismatch with a “young”, “digital” or “agile” identity, rather than as a straightforward capability question.

 

What Transformation Programmes Are Actually Short Of

That instinct is expensive in a way that goes beyond tribunal awards. Transformation programmes run into trouble for reasons that have nothing to do with technical skill: unclear governance, resistance treated as a communications problem rather than early diagnostic information, decisions made by people who have never been accountable for the outcome. Institutional knowledge, stakeholder trust built over years, and the judgement to recognise when a plan won’t survive contact with how the organisation actually operates are not junior capabilities. Screening for “young and agile” screens that experience out at precisely the point a programme needs it most, and does so before anyone has assessed whether the person applying could actually do the job.

 

The Fix Was Never Complicated

EasyJet’s fix did not require lowering a bar. It named the specific bias, redesigned recruitment and onboarding around it, then published the retention data alongside the recruitment numbers rather than stopping at the headline. Technology employers already have research going back years, and a tribunal ruling now sitting on the public record, telling them the same bias exists inside their own hiring and promotion decisions.

 

The Question Worth Asking Before the Next Senior Hire

What’s missing isn’t evidence. It’s a leadership team willing to treat this as a workforce design problem rather than a hiring afterthought. Before the next transformation lead, architect, or programme director role goes out with language built around “energy” or “digital native” instincts, it is worth asking what specific capability that language is actually screening for, and whether the organisation can afford to keep losing the experience it screens out along with it.