Urgent.News

What's breaking now, across thousands of outlets.

AI

Your LLM Types One Token at a Time. It Doesn't Have To.

Every token your LLM emits costs one full forward pass through the entire model. Seventy billion parameters loaded from memory, multiplied, discarded — for a single token. Then again. And again. This is why the big models feel slow, and it's the single most expensive habit in production inference. Speculative decoding breaks the habit. Draft a handful of tokens with something cheap, verify them…

Every token emitted by an LLM costs the full forward pass through the entire model, which is why large models feel slow. Speculative decoding breaks this habit by drafting a few tokens with a cheap model, verifying them all in one pass of the big model, and keeping only the accepted tokens. This can achieve the same output distribution at 2-3x the speed.

In 2026, this is becoming the default inference method. Speculative decoding is like writing a test answer, then checking all the words at once instead of one at a time. A draft model proposes K tokens, and the big target model checks all K in one pass. The key is the acceptance rule: the target accepts a drafted token with probability min(1, q(d)/p(d)) based on how likely it thinks the token is.

Bad drafts are rejected more often and don't corrupt the output. The two key numbers are α (the acceptance rate) and N (the draft length). With α = 0.8 and N = 5, you expect ~3.4 tokens per target forward pass, achieving roughly 3x speedup. Improving α is the main goal for speeding up inference. As methods advance, draft models become more accurate, leading to even greater speedups.

However, production results show that speedup depends on batch size and concurrency. EAGLE-3, the current state-of-the-art, achieves 2-3x speedup over vanilla decoding, but there are diminishing returns at higher batch sizes. Speculative decoding is an optimization technique that converts idle GPU compute into free tokens, but the effectiveness varies based on the specific model and workload.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

An LLM observability platform stores prompts, and prompts are the application

An LLM observability platform stores prompts, and prompts are the application A title query for Langfuse returns 346 matches in ZoomEye. The number is small and the contents are unusual.

  • LLM observability platform records prompts as application logic
  • Traces contain prompts, model parameters, completion details
  • Langfuse indexes sensitive traces, including potential secrets

Tests are the only fixed point left when the implementation is disposable

Originally published at https://aicoding-guide.com . The more disposable the implementation becomes, the more the spec moves into the tests. That is the claim.

  • Tests serve as the fixed point in disposable implementations.
  • Tests guide correctness when code is rewritten repeatedly.
  • Certain domains require alternative fixed points beyond tests.

The agent finished. Who turns off the VM?

Suppose a coding agent opens a pull request at 6 p.m. The tests pass, a preview is running, and the reviewer has already logged off. The agent's task is complete.

  • Agent session ends when task result is saved
  • Preview retained only until explicit expiry time
  • Workspace removal depends on commit status

More from Thursday 8 October →