Story
arxiv_cs_lg ยท Oct 6, 2026 ยท paper
arxiv.orgOct 6, 2026
original source linked
In brief
Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model, referred to as the \emph{draft model}, to generate candidate tokens...