{
  "id": 2302498,
  "title": "GitHub traces 7-hour outage to critical infrastructure failure: Here’s what we know",
  "url": "https://urgent.news/2026/08/21/github-traces-7-hour-outage-to-critical-infrastructure-failure-heres",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-08-21T04:10:56.000Z",
  "source": {
    "name": "The Indian Express",
    "slug": "the-indian-express",
    "url": "https://indianexpress.com/article/technology/tech-news-technology/github-7-hour-outage-what-we-know-10842789/"
  },
  "original_language": "en",
  "account": "GitHub experienced a seven-hour and 47 minute outage on Monday, August 17. The platform's blog post revealed that the root cause was a critical infrastructure component located in one of its US data centers. This component failed to scale to accommodate a surge in traffic, leading to capacity pressure that spread throughout GitHub's systems. As a result, users faced authentication issues and disruptions to numerous GitHub services, including github.com, authentication, GitHub Actions, APIs, pull requests, issues, and Copilot. Most services were back up by the end of the day, but some Copilot services took longer to recover.\n\nGitHub took immediate action during recovery, such as rerouting Teams traffic, isolating affected infrastructure, and restoring services in stages. They excluded any code or configuration changes as the cause of the outage, clarifying that the incidents were both capacity failures. Since April, monthly commits on GitHub have increased significantly, from 1.4 billion to 2.9 billion, which contributed to the strain on their systems.\n\nSince the two major outages in August, GitHub has taken steps to prevent similar occurrences. These include applying consistent retry limits, retry budgets, and variable timeouts across service interactions to avoid retry storms and cascading load. They are also reviewing lower-priority CPU and memory alerts to identify potential failure points in the system.\n\nLooking toward the future, GitHub has set three long-term priorities to prevent such outages. These include adding capacity, improving efficiency, and removing architectural bottlenecks. They have also enhanced their infrastructure with over 3 million additional CPU cores, 120 petabytes of high-speed storage, and increased network capacity. Furthermore, GitHub has accelerated its migration to Microsoft Azure, which currently handles roughly 58% of the platform's load and half of all Git operations. The company aims to create an architecture that can scale read capacity linearly with the number of readers, enabling unlimited reads, starting with the largest monorepos. They are also isolating critical systems and eliminating shared dependencies to minimize the impact and likelihood of future outages.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}