Urgent.News

What's breaking now, across thousands of outlets.

More in Tech

That Time Client Retries Turned a Recovery Into a 7-Hour Outage

GitHub had an interesting incident last August. A component in Central US failed under load, and when it started recovering, the recovery took much longer than it should have.

  • Clients repeatedly retried connections after brief outage, causing flood of requests
  • System partially recovered, but retries flooded it again, creating loop
  • Recommended solution: exponential backoff with jitter to spread out retry load

More from Tuesday 25 August →