Pre-Mortem: The Big Four’s AI Citation Problem

On 28 July 2026, PwC Middle East responded to an investigation into four of its own published reports. The investigation, run by the AI-detection company GPTZero, had found fabricated citations, non-existent sources, and, in one report, a teenage blogger with 280 followers cited as an authority on a JPMorgan initiative. PwC’s statement: the company “takes the accuracy of our published research seriously” and was “updating a limited number of supporting citations.”

PwC was not first. It was the fourth.

This is the sixteenth piece in the Pre-Mortem series. Five questions, applied to the public record, before the outcome is known.

 

The Bet

Deloitte, EY, KPMG and PwC are betting that a pattern spanning five publicly documented reports, four countries, and under two years can be absorbed as unconnected incidents rather than treated as a shared problem with a shared cause. Each firm has responded on its own terms: a partial refund from Deloitte, quiet withdrawals from EY and KPMG, a promise to update “a limited number” of citations from PwC. None has published a shared verification standard. None has described what changes in how AI-assisted work is reviewed before the next report carries its name. The bet is that four reputations, built over more than a century, can absorb five independently verified failures of the most basic check a research report is supposed to pass: that the sources it cites exist.

 

The Assumption

Every one of the four firms has offered a version of the same explanation once caught. KPMG cited guidelines requiring human oversight to validate content and verify sources. PwC cited quality control processes it expects all its people to adhere to. The assumption underneath both statements: that a written guideline is itself a control, that if a policy exists, a human somewhere is presumed to have applied it before publication. EY’s report, “Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems,” carried the names of two partners and a senior manager in its byline. GPTZero’s analysis put the document at roughly 72 per cent AI-generated content, with more than half its 27 sources failing to correspond to anything real. Two partners and a senior manager reviewed that document before it went out, in name. What “reviewed” required in practice is the question none of the four firms has answered.

 

The Sequence

December 2024. PwC Middle East publishes “Agentic AI: The New Frontier in GenAI,” later found by GPTZero to contain fabricated citations.

October 2025. KPMG publishes “Total Experience: Redefining Excellence in the Age of Agentic AI.” GPTZero later finds 45 citations, 5 accurate, at least 16 fabricated.

October 2025. Deloitte refunds AU$97,000 of its A$440,000 contract with Australia’s Department of Employment and Workplace Relations, after a fabricated Federal Court quote and references to non-existent research papers are identified.

November 2025. Newfoundland and Labrador’s C$1.6 million Deloitte health workforce report is found to contain fabricated citations, including one crediting a Dalhousie University researcher as author of a paper that does not exist. Premier Tony Wakeham calls it “concerning.” Deloitte stands by its findings.

27 April 2026. South Africa’s draft National AI Policy is withdrawn 17 days after publication, after 6 of 67 citations are found fabricated. Minister Solly Malatsi calls it “an unacceptable lapse.”

14 May 2026. EY withdraws “Points of Attack” after GPTZero finds more than half its 27 sources do not correspond to real material.

12 June 2026. GPTZero publishes its investigation into KPMG. Five days later, this series covers a separate KPMG story without connecting the two.

28 July 2026. GPTZero publishes its investigation into four PwC Middle East reports. PwC responds that it is updating “a limited number of supporting citations.”

 

 

The Pager

Five public failures, four countries. Three were identified by the same three researchers, Paul Esau, Om Ogale and Alex Cui, working at GPTZero, not at any of the firms and not at any client who paid for the work. Every firm-level response has stopped at the firm: a refund, a report removed from a website, a statement that guidelines exist. No named individual at any firm has been identified as responsible for approving a document whose sources were not real. The one structural change on record did not come from a firm. Newfoundland and Labrador overhauled its own procurement process, requiring disclosure of AI use in future contracts. The government fixed what the contractor did not.

 

The Proof

None of the four firms has published a verification standard: a description of what checking a citation actually involves before a report carries its name. That is the proof measure, not an apology and not a quiet correction, but a public description of the review step, specific enough to be checked against the next report. The IAASB, the International Auditing and Assurance Standards Board, is revising ISA 500, the international standard governing what constitutes sufficient, appropriate audit evidence. That project is still at the research stage and covers formal audit engagements, not the thought-leadership publishing where three of these five failures occurred. Until one firm publishes what verification looks like in practice, every new report each of them publishes resets the same test.

 

Verdict

If one firm publishes a specific, checkable verification standard before a sixth incident surfaces, it becomes the reference point the other three are measured against, the position peer accountability once created around data breach disclosure, where one actor’s transparency made silence from the others harder to sustain. Newfoundland’s government has already shown the structural fix is available: a procurement clause requiring AI disclosure, written in days. If no firm moves first and a sixth incident surfaces, the pattern stops reading as isolated mistakes and starts reading as an industry’s operating baseline. Five failures in under two years, three caught by the same outside team. The firms selling AI governance advisory to clients have not yet demonstrated they can apply the same standard to their own published work. The next report each of them publishes is the test.