Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding - NVIDIA Developer
NVIDIA Developer publishes the DFlash speculative-decoding technique, claiming up to 15x inference throughput gains on Blackwell GPUs.
3 items · 3 sources · 3 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
vLLM and AMD Quark shipped EAGLE-3 speculative decoding for AMD Instinct GPUs on Jul 13, reporting up to 2.00x throughput for Kimi-K2.5 and 1.79x for MiniMax-M2.5 — the first non-NVIDIA vendor to validate speculative decoding as a production latency lever.
Speculative decoding — drafting candidate tokens ahead of the target model to cut serving latency — is a known lever for LLM inference cost. NVIDIA's June 23 developer post introduced DFlash, a speculator built on the target model's own KV projections, claiming up to 15x throughput gains on Blackwell GPUs, and Modal adopted it in production days later with its own open-source speculator models.
State over time
NVIDIA Developer publishes the DFlash speculative-decoding technique, claiming up to 15x inference throughput gains on Blackwell GPUs.
Modal names DFlash as the technique behind its own state-of-the-art latency results, releasing open-source speculator models built with Z Lab and SGLang.
AMD and vLLM ship EAGLE-3 speculative decoding tuned with AMD Quark for Instinct GPUs, the first non-NVIDIA vendor validation of the pattern.
NVIDIA Developer publishes the DFlash speculative-decoding technique, claiming up to 15x inference throughput gains on Blackwell GPUs.
Modal names DFlash as the technique behind its own state-of-the-art latency results, releasing open-source speculator models built with Z Lab and SGLang.
AMD and vLLM ship EAGLE-3 speculative decoding tuned with AMD Quark for Instinct GPUs, the first non-NVIDIA vendor validation of the pattern.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.