{"slug":"dflash-speculative-decoding-on-nvidia-blackwell","label":"Speculative Decoding","item_count":3,"day_count":3,"source_count":3,"first_seen":"2026-06-23T15:14:05+00:00","last_updated":"2026-07-13T00:00:00+00:00","via_scout":true,"generated_at":"2026-07-14T15:04:00.047823+00:00","sources":["modal_blog","search_llm_ops_news","vllm_blog"],"days":[{"date":"2026-06-23","items":[{"title":"Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding - NVIDIA Developer","url":"https://news.google.com/rss/articles/CBMixAFBVV95cUxOZzFLRVlRQk80eUVvMEFHOXhvYjBhbmRRdTFFdFFPcmJ1cVZ2aW5wc3J1RkxqLXpvUEgzaUlyblp2amNNdGR0VzhuZ3pjay1mUW1ZNGdaX1BQZjVhdnp5Qjh3M3Q0amhoUUZJaUNpOF9NcjRIZmw2ckFydUVQZDlHekYzSXdIZ0NRaDN6aGVJdWwtYl9FUGp2WDB0Q0F3YnlENi1YODJYdFNWTmg2LUZwSWFIOVhDT3JPZ2x1bDNXUnBWMkdk?oc=5","source":"search_llm_ops_news","type":"news","summary_1line":"Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding NVIDIA Developer","sid":"99bd515fd5fd8083","published":"2026-06-23T15:14:05+00:00","editor_note":"NVIDIA Developer publishes the DFlash speculative-decoding technique, claiming up to 15x inference throughput gains on Blackwell GPUs."}]},{"date":"2026-06-24","items":[{"title":"Achieve state-of-the-art inference latencies with speculative decoding","url":"https://modal.com/blog/achieve-sota-specdec","source":"modal_blog","type":"news","summary_1line":"How Modal and Decagon worked together to cut inference latency - and you can too.","sid":"62173e9d865bdec2","published":"2026-06-24T00:00:00+00:00","editor_note":"Modal names DFlash as the technique behind its own state-of-the-art latency results, releasing open-source speculator models built with Z Lab and SGLang."}]},{"date":"2026-07-13","items":[{"title":"EAGLE-3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD Quark","url":"https://vllm.ai/blog/2026-07-13-eagle-3-amd-instinct","source":"vllm_blog","type":"news","summary_1line":"How AMD Quark trains, quantizes, and serves EAGLE-3 speculative-decoding drafts with vLLM on AMD Instinct GPUs, delivering up to 2.00x throughput gains for Kimi-K2.5 and 1.79x for MiniMax-M2.5.","sid":"f0c08e4beff850db","published":"2026-07-13T00:00:00+00:00","editor_note":"AMD and vLLM ship EAGLE-3 speculative decoding tuned with AMD Quark for Instinct GPUs, the first non-NVIDIA vendor validation of the pattern."}]}],"editorial":{"tldr":"Speculative decoding — drafting candidate tokens ahead of the target model to cut serving latency — is a known lever for LLM inference cost. NVIDIA's June 23 developer post introduced DFlash, a speculator built on the target model's own KV projections, claiming up to 15x throughput gains on Blackwell GPUs, and Modal adopted it in production days later with its own open-source speculator models.","stale":false,"whats_new":"vLLM and AMD Quark shipped EAGLE-3 speculative decoding for AMD Instinct GPUs on Jul 13, reporting up to 2.00x throughput for Kimi-K2.5 and 1.79x for MiniMax-M2.5 — the first non-NVIDIA vendor to validate speculative decoding as a production latency lever.","why_it_matters":"Speculative decoding is no longer an NVIDIA/SGLang-only trick: if you're serving on AMD Instinct, vLLM plus AMD Quark now gives you an off-the-shelf EAGLE-3 path alongside DFlash's Blackwell/SGLang path, so hardware choice no longer forces you to build custom draft models.","take_for_builders":"Running inference on Blackwell with SGLang? Evaluate the open-source DFlash speculators. On AMD Instinct? vLLM plus AMD Quark now ships EAGLE-3 with reported 1.8–2x throughput gains for Kimi-K2.5/MiniMax-M2.5 — check either path before building custom speculative decoding.","status":{"state":"Spreading across hardware vendors","tone":"rising","changed":"2026-07-13","detail":"AMD Instinct gains its own speculative-decoding path (EAGLE-3 via vLLM + AMD Quark), following NVIDIA's DFlash launch and Modal's production adoption on Blackwell.","track":[{"label":"launch","detail":"Jun 23","tone":"launch","weight":25},{"label":"adopted (NVIDIA)","detail":"Jun 24 → Jul 12","tone":"rising","weight":45},{"label":"second vendor (AMD)","detail":"Jul 13 → now","tone":"now","weight":30}]},"beats":[{"kicker":"LAUNCH","tone":"launch","headline":"NVIDIA introduces DFlash, claiming up to 15x inference gains on Blackwell","summary":"DFlash generates draft tokens in parallel using the target model's own KV projections, invented by Jian Chen and collaborators at Z Lab.","sids":["99bd515fd5fd8083"]},{"kicker":"ADOPTED","tone":"rising","headline":"Modal ships open-source DFlash speculator models with Z Lab and SGLang","summary":"Modal reports DFlash as its best-performing speculative-decoding technique, with a Qwen speculator improving over multi-token prediction by 50%+.","sids":["62173e9d865bdec2"]},{"kicker":"NOW","tone":"now","headline":"AMD Instinct gets its own speculative-decoding path via EAGLE-3 and vLLM","summary":"AMD Quark trains and quantizes EAGLE-3 draft models served through vLLM on AMD Instinct GPUs, delivering up to 2.00x throughput for Kimi-K2.5 and 1.79x for MiniMax-M2.5.","sids":["f0c08e4beff850db"]}],"open_questions":["Does DFlash's speedup hold on non-Blackwell GPUs, or is it Blackwell-specific?","Will DFlash itself get ported to AMD/vLLM, or does EAGLE-3 stay the AMD-side technique?"],"provenance":{"99bd515fd5fd8083":{"surfaced_by":"scout"},"62173e9d865bdec2":{"surfaced_by":"scout"}},"generated_at":"2026-07-14T15:03:32Z"}}