Urgent.News

What's breaking now, across thousands of outlets.

Tech

The 12 kubectl Queries We Run Before Every Cloud Cost Review

Originally published on the Professional IT Services blog . Twelve queries show what a Kubernetes cluster costs and where it wastes money: nodes and taints, requests and limits, real usage, top pods, QoS classes, idle requests, pods over their requests, volumes, orphans, leftovers, load balancers and restarts. They need only kubectl , jq and the hcloud CLI, and the output below comes from our…

Before conducting a cloud cost review for a Kubernetes cluster, it's essential to verify several key metrics using a series of 12 specific kubectl queries. These queries help identify potential inefficiencies and areas where costs could be reduced. Conducting these checks regularly is crucial because cluster usage and billing numbers can change over time, rather than relying on a single snapshot.

First, the inventory of nodes is checked, listing their names, types, architectures, CPU capacities, and any taints. One notable finding is that the master node lacks an instance-type label, which affects the price mapping. The second query provides details on CPU and memory requests versus limits for each node, revealing significant disparities in usage.

For example, one node has a CPU limit of 124% of its capacity, while another node shows 204% limit. Thirdly, the cluster's actual CPU and memory usage is assessed using `kubectl top nodes`, highlighting that only 4.7% of CPU is being used on average, with one node at just 14%.

Moving on, the fourth query evaluates which pods are consuming the most memory, identifying a particular database pod using 4,290 MiB. The fifth query categorizes pods based on their Quality of Service (QoS) classes, with 77 pods marked as Burstable and 21 labeled as BestEffort. The sixth query compares the requested resources against the actual usage for each pod, revealing 2.7 GiB of idle memory across two database replicas.

The seventh query identifies pods that are using more resources than their requests, indicating potential inefficiencies. Eight pods exceed their allocated resources. Next, the ninth query examines volume capacities and usage, showing that 14 out of 16 volumes under 10 GB are underutilized, with many holding less than 1% of their capacity.

The tenth query looks for orphaned volumes that are not attached to any pod but still incur costs. No such volumes were found. The eleventh query checks for LoadBalancer and NodePort services that might be billed outside the cluster, identifying one LoadBalancer and one Service. Lastly, the twelfth query checks for any restarts or Out of Memory (OOM) kills, currently reporting no such issues.

Running these 12 queries before every review is vital because they provide a real-time snapshot of the cluster's resource consumption and billing structure. In May 2026, the cluster successfully reduced its CPU limits from 105% to a much more sustainable 77.7%, resulting in a 29% decrease in monthly compute costs from €114.95 to €81.95.

However, just five months later, the CPU limits had increased to 124%, highlighting the need for ongoing monitoring and adjustment. These regular checks ensure that the cluster remains optimized and cost-effective as conditions change.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Your webhook handler was slow, so Shopify sent it again

The server reboots at two on a Saturday morning for a security update, which is what it's supposed to do. The app doesn't come back, because nobody told the process manager to start it on boot.

  • Shopify resends webhooks if app response exceeds five-second limit
  • Duplicate records occur when server takes too long to acknowledge
  • Self-hosted apps at risk due to downtime and lack of resilience

Hey all..

I’m Darshan, a software engineer, founder and builder from India. I spend most of my time building products, experimenting with AI, automating things.

More from Tuesday 6 October →