rm -rf on the wrong server: how GitLab lost its production database
On January 31, 2017, a GitLab engineer ran rm -rf on the PostgreSQL data directory of what he thought was the replica. It was the primary. About 300 GB of GitLab.com's production database was gone within a second or two, and when the team reached for backups, none of the five they had worked. GitLab.com was down for about 18 hours and lost six hours of data for good. It is the most famous…
On January 31, 2017, a GitLab engineer mistakenly ran the command "rm -Rvf /var/opt/gitlab/postgresql/data" on the primary database server instead of the replica server. This command deleted around 300 GB of the production database within seconds, rendering all five backup methods ineffective. The incident caused GitLab.com to be unavailable for approximately 18 hours, with permanent loss of six hours' worth of data.
The five backup strategies employed by GitLab failed either due to configuration issues or unnoticed failures.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.