Tokens Per Watt: Why Your Context Window Is a Power Decision
On an H100, tokens per watt drops 12x between 4K and 64K context. Agents live at the fat end of that curve. The fix comes from semiconductor architecture.
In May 2004, Intel canceled two chip projects, Tejas and Jayhawk, due to cooling issues. Today, data center electricity consumption has surpassed 447 TWh from 2025 levels, with AI-optimized servers accounting for 31% of the power demand in 2026. Gartner predicts power consumption to increase to 132 GW by this year and reach 290 GW by 2030. TSMC emphasizes energy efficiency, targeting a 30% efficiency improvement with each generation, aiming to ship 1,000W chips and potentially develop megawatt-class systems in the future.
The "tokens per watt" metric becomes crucial when the primary constraint shifts from "how much compute can I buy" to "how much power can I plug in." OpenAI's Jalapeño ASIC, disclosed at Hot Chips, delivers 700W compared to NVIDIA's equivalents drawing 1,200 to 1,400W. This watt budget dictated the architecture. The study "The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency" reveals that when context length doubles, tokens per watt halves.
For instance, serving 2K concurrent sequences on H100-SXM5 with 8 TP and fp16 results in 598 W at 35.0 tok/W, while 16K concurrently requires 557 W at 4.69 tok/W. The GPU's power consumption remains relatively constant, while the number of concurrent sequences decreases linearly with increasing per-token memory footprint. The spread in performance is close to 40x across the 2K–128K range.
The GPU burns roughly the same watts holding 512 sequences as holding 8. However, the numerator collapses as the fixed KV-cache budget accommodates fewer concurrent sequences due to growing per-token memory footprint, leading to linearly decreasing throughput while power stays flat. A significant portion of the power consumption at higher context lengths is just the cost of being powered on, with H100 idle power measured at 300W, accounting for about 43% of TDP.
Jalapeño's NUMA architecture pairs each core slice with a dedicated HBM4 slice to keep the KV-cache local. Despite this, it still draws ~550W sustained, indicating a significant energy requirement even for lightly loaded long-context pools. OpenAI's design for interactive agents emphasizes the importance of this energy-intensive requirement.
Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.