A hung test held the only runner allowed to deploy our hotfix
In August a configuration change made our checkout reject one card type, and the fix was a single line. It was merged at twenty past two. It reached production at twenty to six, and for almost all of the time in between it sat in a queue with a status of queued, which reads as if something is about to happen. Production deploys run on a self hosted runner inside our network, labelled deploy,…
In August, a configuration change caused the checkout to reject one card type. The fix was a single line of code, merged at 2:02 PM. The change initially sat in a queue with a status of "queued," indicating an upcoming action. Production deploys were executed on a self-hosted runner, labeled "deploy," as it was the only machine capable of reaching the cluster API.
A year prior, the integration test job was assigned the same label because it required a database residing on the same private network. Consequently, the machine designated for deployment also became the host for the slowest tests. During the afternoon, an integration test opened a connection to a message broker in a test environment that was undergoing reconstruction.
The message broker responded to the connection but did not provide an appropriate reply. By default, the client library had no read timeout, the test framework lacked a per-test limit configuration, and the job inherited the platform's default six-hour timeout. The test remained in a blocking read state, the runner displayed a "busy" status, and the hotfix was delayed behind it.
The incident was discovered at 5:30 PM when someone noticed the runner's actual activity instead of focusing on the hotfix run. The deployment concluded nine minutes after canceling the test. Since then, several changes have been implemented. Every job now includes a timeout-minutes setting, which is approximately three times the normal duration, derived from the last month's run metrics.
A lint step now rejects any workflow that fails to include this timeout parameter. Additionally, the test framework incorporates a per-test timeout, causing hangs to fail specific tests rather than consuming the entire job's resources. The deploy runners are now dedicated solely to deployment tasks, with two instances located in different zones.
Tests requiring access to the private network run on dedicated pools with separate labels. Moreover, any job queued on the "deploy" label for more than five minutes triggers an alert to the on-call personnel, as a waiting deployment can become a critical aspect of an incident. The decision to set the default timeout to six hours was not arbitrary; it represents the maximum duration the platform will wait before terminating a connection, serving as a ceiling rather than a predetermined budget.
This particular scenario involved the deployment machine standing between a malfunctioning checkout and the critical fix that ultimately resolved the issue. The incident was reported by Sergey Shinder.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.