The September AI Outage Had Two Real Causes. Almost Everyone Reported One.

On the morning of 3 September 2026, ChatGPT, Claude and Grok all failed within roughly ninety minutes of each other. Within hours, most coverage had settled on a single explanation: a shared Microsoft Azure failure had taken down three competing AI platforms at once. The postmortems that actually exist tell a different story, and the difference matters more than the outage itself.

 

What Each Company Actually Said

SpaceXAI was the most direct. The company confirmed a failure at its Memphis compute facility, apologising to what it called its “impacted compute partners,” language that implies other organisations lease capacity in the same facility. Grok was down for roughly three and a half hours. OpenAI gave its own, separate explanation: an engineer identifying as the incident commander stated on Hacker News that the cause was a routing error inside OpenAI’s own infrastructure, surfacing roughly ninety minutes after SpaceXAI’s Memphis failure and unconnected to the GPT-6 Astra launch that followed the next day. Anthropic never published a technical cause at all. It confirmed elevated errors across several Claude models starting around 9:26am ET and reported the issue resolved, but the actual mechanism behind it was never disclosed.

 

Why the Azure Theory Doesn’t Hold Up

The Azure theory spread fast because it was the simplest explanation available in the first few hours, and it fit an existing assumption that most enterprise AI risk registers already carry: if these vendors compete, their infrastructure probably doesn’t overlap. It did not hold up. Microsoft logged no Azure incident for that window, and Cloudflare said in a statement that it was not experiencing any significant service disruptions. The timeline argues against one shared trigger too, SpaceXAI’s and Anthropic’s incidents began four minutes apart from each other, and OpenAI’s followed roughly ninety minutes later, a pattern that fits three separate incidents more comfortably than one shared one.

 

The Question Nobody Answered

The part worth noticing is not that the Azure story was wrong. It is that the two companies which did give an explanation gave two different ones, and the one that stayed quiet is the one whose incident opened four minutes before SpaceXAI’s, the tightest timing overlap of the three. Nobody, including Anthropic, ever confirmed or ruled out whether Claude’s infrastructure touches the same Memphis facility SpaceXAI apologised to its partners about. A vendor risk register cannot verify a question its own vendor never answers.

None of this means multi-vendor AI strategy is worthless. It means the diversification most risk registers assume is rarely something a customer can actually verify from the outside, and the one moment it gets tested, a real outage, is also the moment vendors are least likely to disclose the detail that would let you check. Two of three companies gave a specific, different cause. One gave none. Whichever of those a programme is depending on, the honest entry in a vendor risk register reads “unconfirmed,” not “diversified.”