LLM Digest
Subscribe

AI Storyline

3 items · 3 sources · 3 days

View as JSON

Operational story trace

Speculative Decoding

Current stateSpreading across hardware vendorsstatus changed Jul 13

Latest change

vLLM and AMD Quark shipped EAGLE-3 speculative decoding for AMD Instinct GPUs on Jul 13, reporting up to 2.00x throughput for Kimi-K2.5 and 1.79x for MiniMax-M2.5 — the first non-NVIDIA vendor to validate speculative decoding as a production latency lever.

Earlier contextThe story so far

Speculative decoding — drafting candidate tokens ahead of the target model to cut serving latency — is a known lever for LLM inference cost. NVIDIA's June 23 developer post introduced DFlash, a speculator built on the target model's own KV projections, claiming up to 15x throughput gains on Blackwell GPUs, and Modal adopted it in production days later with its own open-source speculator models.

editor-curated · source-linked

State over time

● launch · Jun 23second vendor (AMD) · Jul 13 → now ●
  • launch · Jun 23
  • adopted (NVIDIA) · Jun 24 → Jul 12
  • second vendor (AMD) · Jul 13 → now
LAUNCH · Jun 23
NVIDIA introduces DFlash, claiming up to 15x inference gains on Blackwell
1 source · scout · show source ▾
ADOPTED · Jun 24
Modal ships open-source DFlash speculator models with Z Lab and SGLang
1 source · scout · show source ▾
NOW · Jul 13
AMD Instinct gets its own speculative-decoding path via EAGLE-3 and vLLM
1 source · show source ▾

What to watch — open questions

  • Does DFlash's speedup hold on non-Blackwell GPUs, or is it Blackwell-specific?
  • Will DFlash itself get ported to AMD/vLLM, or does EAGLE-3 stay the AMD-side technique?
How this thread was built
scout surfaced 2editor wrote the arc · 3 beatswatcher 1 status change

Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.