AI Briefing: September 4, 2026 — SpaceXAI's Memphis outage and Anthropic's Colossus 1 lease.

THE MORNING THREE RIVALS WENT DARK

Grok went first. xAI's status page logged a models outage beginning around 6:30 a.m. Pacific on Thursday, September 3, and it stayed broken for roughly three and a half hours. By mid-morning Eastern time, the trouble had spread: Downdetector recorded a flood of reports between 10:30 and 11:00 a.m. ET, ultimately topping 35,000 for ChatGPT, about 1,400 for Claude, and 1,200 for Grok in the U.S. alone. Cursor, the AI coding editor that routes requests through several of these same model providers, reported its own disruption in the same window. Google's Gemini, notably, kept running the whole time — a detail that matters once you start asking why three separate companies broke at once.

WHAT EACH COMPANY'S OWN STATUS PAGE ADMITTED

OpenAI opened an incident for ChatGPT at 10:58 a.m. Pacific and moved it to "monitoring" by 11:50 a.m. after applying a mitigation — roughly fifty minutes of acknowledged trouble, by its own account. Anthropic's status page listed elevated errors across Claude Opus 5, Opus 4.8, and Opus 4.6, along with the newer Mythos and Fable 5 and 5.1 models, and named claude.ai, the Claude API, Claude Code, and Claude Cowork as affected services. xAI's own explanation, once it came, was the most specific of the three: a hardware-level failure at a named physical facility, not a vague "elevated error rate." That specificity is what makes the other two companies' silence about a cause stand out.

AN APOLOGY THAT NAMED A PARTNER WITHOUT NAMING ONE

SpaceXAI, xAI's parent company, posted on X: "We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We'd also like to apologize to our impacted compute partners. All systems have now been restored and are functioning nominally." Elon Musk followed up separately, saying the company was "taking corrective action to ensure this does not happen again." Neither post named which partners those were. But there's only one obvious candidate renting meaningful capacity at that facility — and it's a company xAI competes directly against for frontier-model customers.

THE DEAL THAT MAKES ANTHROPIC ONE OF THOSE PARTNERS

On May 6, 2026, xAI signed Anthropic to what became one of the largest infrastructure contracts in the industry's history: roughly $1.25 billion a month, running through May 2029, for access to essentially the entire capacity of Colossus 1 — xAI's Memphis, Tennessee supercomputer, built around more than 220,000 Nvidia GPUs (a mix of H100, H200, and GB200 accelerators) drawing over 300 megawatts. The arrangement made sense for both sides at the time: xAI had shifted its own frontier training to a newer cluster, Colossus 2, leaving Colossus 1 idle capacity to lease, and Anthropic needed compute badly enough to rent it from a direct rival. Four months later, that rival's data center had a bad morning — and Anthropic's own infrastructure sits inside it.

THE THEORIES NOBODY WOULD CONFIRM

Asked by reporters why ChatGPT and Claude went down the same morning as Grok, neither OpenAI nor Anthropic pointed to an external cause. Other explanations circulated regardless: a possible failure in Microsoft Azure's East US region, which hosts meaningful AI-workload traffic, and a Cloudflare disruption compounding it. Neither held up under scrutiny. Cloudflare stated flatly: "Our services are operating normally, and any reporting that deviates from this is incorrect." AWS, Google Cloud, and Microsoft Azure's own status pages showed no relevant incidents during the window in question. That leaves the Memphis explanation as the only one any company has actually put its name to — and the only one two of the three affected companies have declined to either confirm or rule out.

WHY THIS MATTERS FOR TEAMS BUILDING ON AI

If you route requests to "different" AI providers as a resilience strategy — falling back from Claude to GPT to Grok when one degrades — the assumption underneath that plan is that these are independent systems failing independently. The Colossus 1 lease says otherwise: two labs racing each other on capability can still be sharing a power substation, a cooling system, or a fiber run underneath the competition. Nobody outside these companies knows for certain that Thursday's three-way outage traces to one Memphis facility rather than three unrelated coincidences. But the one company that did name a cause named a facility a second company depends on, and that second company had every opportunity to say "not us" and didn't. Before you count "multiple providers" as a real hedge against downtime, it's worth finding out whether their infrastructure is actually as separate as their marketing suggests.