Urgent.News

What's breaking now, across thousands of outlets.

Tech

A hung test held the only runner allowed to deploy our hotfix

In August a configuration change made our checkout reject one card type, and the fix was a single line. It was merged at twenty past two. It reached production at twenty to six, and for almost all of the time in between it sat in a queue with a status of queued, which reads as if something is about to happen. Production deploys run on a self hosted runner inside our network, labelled deploy,…

In August, a configuration change caused the checkout to reject one card type. The fix was a single line of code, merged at 2:02 PM. The change initially sat in a queue with a status of "queued," indicating an upcoming action. Production deploys were executed on a self-hosted runner, labeled "deploy," as it was the only machine capable of reaching the cluster API.

A year prior, the integration test job was assigned the same label because it required a database residing on the same private network. Consequently, the machine designated for deployment also became the host for the slowest tests. During the afternoon, an integration test opened a connection to a message broker in a test environment that was undergoing reconstruction.

The message broker responded to the connection but did not provide an appropriate reply. By default, the client library had no read timeout, the test framework lacked a per-test limit configuration, and the job inherited the platform's default six-hour timeout. The test remained in a blocking read state, the runner displayed a "busy" status, and the hotfix was delayed behind it.

The incident was discovered at 5:30 PM when someone noticed the runner's actual activity instead of focusing on the hotfix run. The deployment concluded nine minutes after canceling the test. Since then, several changes have been implemented. Every job now includes a timeout-minutes setting, which is approximately three times the normal duration, derived from the last month's run metrics.

A lint step now rejects any workflow that fails to include this timeout parameter. Additionally, the test framework incorporates a per-test timeout, causing hangs to fail specific tests rather than consuming the entire job's resources. The deploy runners are now dedicated solely to deployment tasks, with two instances located in different zones.

Tests requiring access to the private network run on dedicated pools with separate labels. Moreover, any job queued on the "deploy" label for more than five minutes triggers an alert to the on-call personnel, as a waiting deployment can become a critical aspect of an incident. The decision to set the default timeout to six hours was not arbitrary; it represents the maximum duration the platform will wait before terminating a connection, serving as a ceiling rather than a predetermined budget.

This particular scenario involved the deployment machine standing between a malfunctioning checkout and the critical fix that ultimately resolved the issue. The incident was reported by Sergey Shinder.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Our Tickets Closed Themselves Before the Night Shift Came Back

Our service desk tool closes a ticket automatically when the requester has not replied for three working days. It is the supplier's recommended setting, and for people at head office it works well.

  • Service desk tool automatically closes tickets after three working days without response
  • Night shift supervisor's ticket closed prematurely, requiring new ticket creation
  • Timer now counts requester's shifts instead of calendar days to address issue

Your Mentee Has Overtaken You in Something

It usually shows up as a small moment. You are talking about their work, you offer your usual view on how the caching layer should be structured, and they answer politely with something you did not…

  • Mentee now surpasses mentor in specific expertise
  • Mentor acknowledges shift in relationship dynamics
  • Relationship evolves into equal partnership

Custom Web Development vs Templates: A Practical Decision Framework

"Should we build it custom or use a template?" is one of the most common questions in web projects, and it is usually answered with opinions instead of criteria.

  • Templates suitable for brochure sites, blogs, and portfolios
  • Custom development better for complex business logic, integrations, and performance control

Why Most Telegram Store Bots Break at Scale (and How to Fix It)

Telegram is the cheapest storefront on the internet: no app store review, no hosting bill for the storefront itself, and an audience that already lives inside the app.

  • Telegram bots fail at scale due to in-memory state storage.
  • Payment webhook retries cause double deliveries without idempotency keys.
  • Fulfilment intertwined with bot operation leads to slow uploads and user experience issues.

More from Tuesday 29 September →