Urgent.News

600+ sources. One page. See who else covered it.

Editions

More in AI

Speculative Decoding: Faster On-Device LLMs

Speculative decoding makes a large language model generate text faster on a constrained device without changing what it produces.

  • Speculative decoding speeds up on-device LLMs by generating candidate tokens in advance.
  • Small draft model evaluates candidate tokens against larger model's acceptance criteria.
  • Speed-up ranges from 2x to 3x on predictable text, but less effective on high-entropy text.

More from Wednesday 5 August →