Your LLM Types One Token at a Time. It Doesn't Have To.
Every token your LLM emits costs one full forward pass through the entire model. Seventy billion parameters loaded from memory, multiplied, discarded — for a single token. Then again. And again. This is why the big models feel slow, and it's the single most expensive habit in production inference. Speculative decoding breaks the habit. Draft a handful of tokens with something cheap, verify them…
Every token emitted by an LLM costs the full forward pass through the entire model, which is why large models feel slow. Speculative decoding breaks this habit by drafting a few tokens with a cheap model, verifying them all in one pass of the big model, and keeping only the accepted tokens. This can achieve the same output distribution at 2-3x the speed.
In 2026, this is becoming the default inference method. Speculative decoding is like writing a test answer, then checking all the words at once instead of one at a time. A draft model proposes K tokens, and the big target model checks all K in one pass. The key is the acceptance rule: the target accepts a drafted token with probability min(1, q(d)/p(d)) based on how likely it thinks the token is.
Bad drafts are rejected more often and don't corrupt the output. The two key numbers are α (the acceptance rate) and N (the draft length). With α = 0.8 and N = 5, you expect ~3.4 tokens per target forward pass, achieving roughly 3x speedup. Improving α is the main goal for speeding up inference. As methods advance, draft models become more accurate, leading to even greater speedups.
However, production results show that speedup depends on batch size and concurrency. EAGLE-3, the current state-of-the-art, achieves 2-3x speedup over vanilla decoding, but there are diminishing returns at higher batch sizes. Speculative decoding is an optimization technique that converts idle GPU compute into free tokens, but the effectiveness varies based on the specific model and workload.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.