DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
DARTree proposes diffusion-based draft trees that predict a whole token block per step instead of one token at a time.
3 items · 3 sources · 3 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
A new paper (AsymSpec, Aug 26) targets speculative decoding specifically at agentic pipelines, arguing that generic techniques don't address the latency cost of long, tool-call-and-retrieval-inflated context that agents accumulate.
Speculative decoding — verifying multiple draft tokens per step to cut inference latency — is now developing on two separate fronts: new drafting research and production serving guides. A mid-August paper (DARTree) pushed the research side toward diffusion-based draft trees that predict whole token blocks instead of single tokens.
Arc
DARTree proposes diffusion-based draft trees that predict a whole token block per step instead of one token at a time.
vLLM publishes a practical speculative-decoding tuning guide for AMD GPUs covering MTP, EAGLE-3, DFlash, and DSpark.
AsymSpec proposes context-asymmetric speculative decoding aimed specifically at agentic LLM pipelines' inflated context cost.
DARTree proposes diffusion-based draft trees that predict a whole token block per step instead of one token at a time.
vLLM publishes a practical speculative-decoding tuning guide for AMD GPUs covering MTP, EAGLE-3, DFlash, and DSpark.
AsymSpec proposes context-asymmetric speculative decoding aimed specifically at agentic LLM pipelines' inflated context cost.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.