EAGLE-3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD Quark
vLLM's own blog: AMD Quark-trained EAGLE-3 drafts cut serving latency up to 2x for Kimi-K2.5 on AMD Instinct GPUs.
4 items · 2 sources · 4 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
Netflix disclosed the internals of its in-house LLM serving platform, built on Triton and vLLM — the thread's first large-scale production-adoption proof point after three straight vLLM engine upgrades this month.
Across three July posts, vLLM's own blog shipped AMD-backed EAGLE-3 speculative decoding, a TileRT-powered latency-optimized decode path, and an Attention-FFN disaggregation plugin for MoE serving — each landing through the same V1 connector interface without disrupting existing deployments.
Arc
vLLM's own blog: AMD Quark-trained EAGLE-3 drafts cut serving latency up to 2x for Kimi-K2.5 on AMD Instinct GPUs.
vLLM pairs prefill with a separate TileRT decode engine behind one shared serving layer, with no changes needed to existing deployments.
vLLM's AFD plugin disaggregates attention and FFN computation for MoE models, adding GPU and Ascend NPU backend support.
Netflix goes public with its in-house LLM serving platform built on Triton and vLLM — the thread's first large-scale production-adoption signal.
vLLM's own blog: AMD Quark-trained EAGLE-3 drafts cut serving latency up to 2x for Kimi-K2.5 on AMD Instinct GPUs.
vLLM pairs prefill with a separate TileRT decode engine behind one shared serving layer, with no changes needed to existing deployments.
vLLM's AFD plugin disaggregates attention and FFN computation for MoE models, adding GPU and Ascend NPU backend support.
Netflix goes public with its in-house LLM serving platform built on Triton and vLLM — the thread's first large-scale production-adoption signal.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.