A third of our releases went out through the emergency path
A deploy at eleven on a Tuesday night broke search for about forty minutes. It had skipped the integration suite and the staging soak, and it had needed no second approver, all of which was allowed, because it had gone out through the pipeline we built for emergencies. I went to look at how often that happened, expecting to find a handful. Thirty-one of the previous ninety production deploys had…
In a Tuesday night deploy, it was discovered that a third of the releases went out through an emergency path, bypassing the usual integration suite and staging soak. Despite the lack of tests, reviews, and approvals, this emergency path was used for 31 out of 90 production deploys in the previous 90 days. The standard pipeline, which took 52 minutes, was deemed too slow, with 34 of the 31 emergency deploys needing a retry roughly one in five times.
By labeling every deploy with the path it took and sharing the usage split in a weekly team message, usage of the emergency path dropped by half in the first two weeks. The emergency path now requires a typed reason, announces itself in a channel, and runs the full suite after deployment with an automatic rollback if it fails. This change resulted in emergency usage dropping to about two out of every ninety deploys.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.