Three AI Providers Failed in the Same Window. Not One Published a Cause.

· CX Pulse

Anthropic, OpenAI and Grok all failed on September 3. None of them published a cause, and the apology was still yours to write.

Between 8:37am and 1:07pm Eastern yesterday, three of the largest AI providers on the market failed inside the same window. If your customers talk to something that runs on one of them, they hit an error, and you are the one who had to explain it.

The times below come from the providers' own incident feeds rather than from the coverage.

Anthropic opened its first incident at 12:37 UTC for Claude Sonnet 5 and closed it nineteen minutes later. At 13:26 UTC it opened a second one, "Elevated errors for multiple models," which eventually named six: Mythos and Fable 5.1, Mythos and Fable 5, Opus 5, Opus 4.8 and Opus 4.6. That one ran until 16:16 UTC. Just under three hours.

OpenAI opened "Elevated errors across ChatGPT and Codex" at 14:58 UTC and resolved it at 16:55. Grok failed in the same window. xAI's status page sits behind a challenge screen and doesn't serve a public incident feed, so the clearest surviving record of that outage is a downstream one.

One company filed three incident reports in 96 minutes

Cursor runs on models from several providers. Its status page yesterday reads like a stress test nobody scheduled.

Three incidents, three providers, ninety-six minutes. Routing across multiple model vendors is the standard answer to the question of what happens when your AI provider goes down. Yesterday that answer was tested in production, and for at least one company running it at scale, it didn't hold.

Nobody said what broke

At 13:41 UTC Anthropic wrote that it had "identified the cause" and was working on a fix. Six updates later, at resolution, it still hadn't said what that cause was. OpenAI never offered one at all. Its updates ran from "we are investigating the issue for the listed services" to "we have applied the mitigation," with nothing in between. One theory circulated publicly, pointing at Azure. No provider confirmed it.

So the record available to anyone downstream is that it broke, and then it stopped being broken. That's the whole thing.

If you run customer experience at scale

Four items on this belong in writing while yesterday is still fresh.

Service credits are computed off the vendor's own severity grade, not off yours. OpenAI classified its own incident yesterday as minor impact. Cursor, sitting downstream of all three, graded its worst incident of the same afternoon as critical. Both are accurate from where each company sits, and only one of them is the number your contract pays out against.

Your root-cause obligation doesn't inherit your supplier's silence. If you owe a regulated customer or an enterprise account an RCA inside a fixed window, "our provider hasn't told us" is what you file, with your name on it.

Three upstream incidents don't reconcile into one story for your customers. You had one outage. Your vendors have three postmortems, on three clocks, in three formats, none of which reference each other. Whoever writes your customer communication is doing that merge by hand.

Multi-vendor routing sits in a lot of risk registers as the mitigation for exactly this scenario. Yesterday produced real evidence about whether that control holds. It's worth capturing before the memory of it softens. The related question, and the one that got answered informally yesterday by whoever happened to be watching the queue: who is authorized to declare your own incident when no vendor has declared one?

If you run something smaller

The same afternoon looked different and had the opposite problem. A chat widget on a small site showed a spinner and then an error for a couple of hours, and there was no status page of yours for anyone to check. The only person who told your customer that something was wrong was your customer.

A large organization has the incident process and yesterday found it depending on facts a supplier never released. A smaller one has the fact, which is simply that the thing stopped working, and nowhere to put it.

The part to do this week

Decide in advance what you say, and where you say it, when the model behind your front door stops answering and nobody upstream explains why. Write that message now, while nothing is broken, and pick the surface it goes on before you need it.

Recovery yesterday took between two and three and a half hours depending on which provider you were sitting behind. Every minute of that was a minute your customer spent forming an opinion about you, not about them.

Sources: Anthropic status, OpenAI status, Cursor status.