When the cluster looked idle but nothing would schedule
The first thing I checked was the obvious one: CPU. Our non-prod EKS workers are Karpenter-managed, and application pods had been sitting in Pending long enough that Istio and app syncs were backing up. kubectl top nodes looked almost embarrassingly healthy — low CPU, memory in a comfortable band. If you squinted at the dashboard, you'd swear we had room. Karpenter disagreed. Its logs kept…
The team noticed that their non-production Kubernetes cluster, managed through EKS and Karpenter, was experiencing a scheduling issue where pods remained in a pending state. CPU resources appeared to be abundant, yet the pods were unable to be launched. The team examined the CPU and memory usage on the nodes through kubectl top nodes, and they seemed to have ample capacity.
However, Karpenter's logs kept repeating the same error: "Failed to schedule pod, all available instance types exceed limits for nodepool." This discrepancy between the apparent available resources and the inability to schedule new pods led to a frustrating hour of false confidence.
Upon further investigation, the team discovered that the primary app pool, ec2nodepool, was configured to provision Graviton workers through Karpenter. The limits and instance rules for this pool were defined in GitOps using helm-values/karpenter/nodepool.yaml. The node pool was intended to scale on demand in a few Availability Zones, with a preference for ARM64 on-demand instances.
The workloads pending were significant, typically requiring around 700m CPU and 2Gi memory per pod, which was typical for the services being rolled out.
The Kubernetes scheduler places pods based on requests, not the actual live usage as shown by top. Karpenter, when deciding whether to add a node, charged the full EC2 instance shape against the NodePool's limits, not the remaining allocatable resources on existing nodes or the live top usage. This led to a misunderstanding, as Karpenter was not adding nodes even though the scheduler was showing insufficient resources.
The root cause was found to be two-fold. First, the memory limit was incorrectly set as if it were symmetric with CPU, but instance charging was asymmetric for the Graviton c6g-heavy fleet. Each new c6g.large instance required roughly 3.7Gi of counted capacity, while the pool had already consumed approximately 42 CPU and 79Gi of memory against its caps.
No available instance type fit within the remaining memory budget. Second, the node pool was restricted to only two instance types: c6g.large and c6g.xlarge. This limitation prevented Karpenter from launching additional nodes when the memory limit was reached, as there were no other instance types that could accommodate the remaining memory requirement.
To resolve the issue, the team made two changes in GitOps. First, they increased the memory limit from 80Gi to 160Gi, while keeping the CPU limit at 80. This adjustment provided headroom for the CPU budget and allowed more instances to be provisioned. Second, they removed the hardcoded instance type requirement, enabling Karpenter to choose from multiple on-demand ARM64 instance types (r, c, and t) with generation 4 or higher.
This change restored flexibility in instance selection and allowed the node pool to scale appropriately.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.