Urgent.News

What's breaking now, across thousands of outlets.

Tech

Would your Kubernetes cluster survive losing one ESXi host? I built a read-only tool to find out

Kubernetes sees nodes . vSphere sees VMs on hosts . Neither tells you when DRS has put two of your three etcd VMs on the same ESXi host. From then on, one host failure takes your control plane down, and every dashboard still says green. I spent nearly seven years in VMware support, and I couldn't find a tool that joins those two views and simulates host failures against Kubernetes semantics. So I…

Kubernetes clusters can be vulnerable if an ESXi host fails, potentially taking down the control plane. A read-only tool called kube-hostfail was created to simulate host failures against Kubernetes semantics to identify such vulnerabilities. The tool joins Kubernetes node and etcd-member data with vCenter VM and host data, then simulates single and two-host failures.

It checks etcd quorum, API availability, workload capacity, and DRS anti-affinity rules. If any single host failure could take down the cluster, kube-hostfail exits with a code indicating the level of failure. The tool is designed to be installed in a production environment without modifying the existing cluster. It includes various components such as Prometheus, Grafana, Istio, Jaeger, Kiali, and a gateway API, all in new namespaces with sidecars injected only into the demo namespace.

The tool has been verified through unit tests, end-to-end tests, chart and manifest validation, and a full install on a real cluster. However, it has not been verified in a live production environment and does not model datastore/PV placement, network partitions, DaemonSets, or vSphere HA admission control. Users are encouraged to try the tool and report any issues or suggestions for improvement.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Your JSON Parses. So Why Did One Value Disappear?

Imagine reviewing a configuration file for a nightly import job. You expect the job to stop after processing 1,000 rows. But the application keeps going until 10,000.

  • Nightly import job expected to stop at 1,000 rows
  • JSON file had duplicate maxrows key with 1000 and 10000
  • Application retained last value, lost initial 1,000

Edtech Observability Stack — Health Monitoring Through Logs, Metrics, and Result Deadlines

An uptime check can prove that an edtech app answers requests while its nightly roster import produces nothing. The useful trade-off is signal quality versus noise: alert on a missed, overdue result…

  • Monitor edtech app health via structured logs, metrics, and completion events
  • Implement uptime probe to verify promised outcomes like student record imports
  • Use TypeScript implementation with OpenTelemetry metrics for aggregate measurements

Your seed data is lying to you — the case for foreign-key-consistent test data

Here's a bug I've shipped more than once: tests pass locally, demo looks great, then something breaks the moment real relational data shows up.

  • Seed data with foreign keys fails when real relational data introduced.
  • Faker-style tools generate values in isolation, not considering relationships.
  • Automating data generation ensures foreign keys point to existing records.

More from Sunday 11 October →