Secure Speculative Decoding for Large Language Models
Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model, referred to as the \emph{draft model}, to generate candidate tokens and then verifying them with the target model for acceptance or rejection. Prior studies primarily focused on the efficiency-utility trade-off of speculative decoding, e.g., lossy…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.