Urgent.News

What's breaking now, across thousands of outlets.

Tech

GitHub blames 8-hour outage on autoscaling fail and VS Code retry storm

Load balancers buckled after a monitoring blind spot allowed traffic to spiral

GitHub blames 8-hour outage on autoscaling fail and VS Code retry storm

GitHub has detailed the reasons behind its near eight-hour outage that left developers unable to work normally. The incident began at 1328 UTC on August 17 and was not fully resolved until 2115 UTC, affecting multiple services such as Issues, Pull Requests, APIs, Actions, and Copilot. The root cause was a saturated load balancer in the company's Central US facility, exacerbated by a faulty autoscaling policy and a latent retry bug in Visual Studio Code.

An Istio sidecar reaching its concurrency limit triggered the problem, despite autoscaling being expected to add capacity. Optimistic retry logic in VS Code overloaded internal load balancers, amplifying traffic by around 10x and delaying the recovery of the Copilot Token Service. Engineers temporarily reduced gateway retries and configured the load balancers to reject Copilot Token Service requests, but complications like scraping attacks on codeload endpoints delayed the recovery.

GitHub plans to adjust autoscaling policies, review retry limits, audit Istio settings, and address VS Code behavior. This incident may prompt developers to consider alternative solutions, leading to a more bifurcated ecosystem.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theregister.com →

More in Tech

More from Wednesday 19 August →