I Deliberately Destroyed My Kubernetes Cluster at 2 AM. Here's What Died First.
I Deliberately Destroyed My Kubernetes Cluster at 2 AM. Here's What Died First. Chaos engineering is not about breaking things. It's about discovering that your "production-grade" homelab is held together by hope and a single etcd snapshot before someone else finds out for you. The Setup I was lying in bed at 1:47 AM, staring at the ceiling, unable to sleep. Not because of caffeine. Because of a…
At 2 AM, the author deliberately destroyed their Kubernetes cluster to test its resilience. Running a 4-node bare-metal cluster on Talos Linux, Dell OptiPlex control plane, and three Raspberry Pi workers, they employed Chaos Mesh to induce failures. The experiment began with pod chaos, causing random pods to be killed every 30 seconds.
Despite having two replicas, Prometheus lost 15 minutes of data and MarketPulse API returned 500 due to the database being unavailable. Next, a network partition was simulated, isolating one worker and causing Cilium to fail. The Longhorn volume with replicas on both nodes lost quorum, degrading write operations and causing Prometheus to mark pods as DOWN.
The experiment highlighted the need for better handling of stateful workloads and the importance of robust failure recovery in a homelab environment.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.