Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

Tech

GitHub blames 8-hour outage on autoscaling fail and VS Code retry storm

Load balancers buckled after a monitoring blind spot allowed traffic to spiral

GitHub blames 8-hour outage on autoscaling fail and VS Code retry storm

GitHub announced this week's prolonged outage, lasting nearly eight hours, from August 17 at 1328 UTC to August 17 at 2115 UTC. The primary issue stemmed from overloaded load balancers in their Central US data center, which occurred when an Istio sidecar hit its concurrency limit. Autoscaling did not alleviate the problem, as a misconfigured policy monitored the host service but not the sidecar's concurrency limit.

This oversight allowed a failure to escalate. The situation worsened due to optimistic retry logic in Visual Studio Code, which overloaded internal load balancers and amplified traffic by about ten times. Engineers temporarily alleviated the issue by reducing gateway retries and configuring load balancers to reject Copilot Token Service requests with HTTP 403 responses.

Despite most services recovering by 1636 UTC and Actions by 1803 UTC, the Copilot Token Service remained problematic until 2102 UTC. The outage was further exacerbated by scraping attacks on codeload endpoints. GitHub pledged to review and adjust autoscaling policies, retry limits, and Istio concurrency settings, as well as address the VS Code retry bug contributing to the issue.

This incident may prompt some developers to explore alternative solutions, as competitors like Cursor and OpenAI are already gaining traction. GitHub's reliability issues, not new, could lead to a more fragmented developer ecosystem.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in Tech

More from Wednesday 19 August →