ChatGPT, Claude, and Grok all went down at once; enterprises need a backup plan
Enterprises are facing a disturbing new question in the age of AI: What happens when agentic assistants go dark? This became a very real scenario on Thursday, as OpenAI’s ChatGPT, Anthropic’s Claude, and SpaceXAI’s Grok near-simultaneously, and somewhat mysteriously, experienced significant, prolonged outages. Beginning in the morning, Eastern time, several ChatGPT models went down over a roughly…
Enterprises are grappling with a new concern in the age of artificial intelligence: What occurs when AI agents cease functioning? This became a tangible reality on Thursday, as OpenAI's ChatGPT, Anthropic's Claude, and SpaceXAI's Grok all suffered significant, prolonged outages. The outages occurred simultaneously and unexpectedly, causing widespread concern among users and IT teams.
On Thursday morning, several ChatGPT models went down for approximately two hours, while Claude models were inaccessible for four hours, and Grok models were offline for nearly three and a half hours. All three companies acknowledged the issues and implemented fixes. As users grumbled in forums and IT teams worked to restore service, the incident highlighted how hastily organizations have adopted generative AI workflows without considering the potential impact of widespread outages.
AI agents are increasingly handling automated and broader-scale workflows, and companies could find themselves "uncomfortably exposed" if AI suddenly stops functioning, according to technology analyst and journalist Carmi Levy. This incident should serve as a wake-up call for IT leaders who have largely overlooked the costs of AI outages, as the risk is no longer hypothetical.
The outages affected core services, including ChatGPT's search, file uploads, agents, GPTs, voice mode, image generation, and various connectors and apps. OpenAI's Codex services, such as the web, API, command-line interface (CLI), and VS code extension, were also impacted. Claude and Grok experienced similar issues, with some services returning to "healthy" traffic later in the day.
This simultaneous disruption raises concerns about the potential impact on organizations relying heavily on AI platforms. AI agents are becoming critical components of workflows, and extended outages could leave employees stranded without manual alternatives. Levy emphasizes the need for organizations to reassess their disaster recovery and business continuity plans, as well as to train employees to maintain manual skills and be prepared for service interruptions.
According to Info-Tech Research Group's Brian Jackson, a modular architecture for large language models (LLMs) could provide a solution, allowing enterprises to quickly switch to alternative providers if their primary choice experiences a concurrent outage.
Written by urgent.news from Computerworld's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.