LLM Digest
Subscribe

AI Storyline

4 items · 2 sources · 4 days

View as JSON

Operational story trace

Serving Vllm

Latest change

Netflix disclosed the internals of its in-house LLM serving platform, built on Triton and vLLM — the thread's first large-scale production-adoption proof point after three straight vLLM engine upgrades this month.

Earlier contextThe story so far

Across three July posts, vLLM's own blog shipped AMD-backed EAGLE-3 speculative decoding, a TileRT-powered latency-optimized decode path, and an Attention-FFN disaggregation plugin for MoE serving — each landing through the same V1 connector interface without disrupting existing deployments.

editor-curated · source-linked

Arc

Jul 13Jul 27 · now
SPECULATIVE DECODING · Jul 13
vLLM ships EAGLE-3 speculative decoding for AMD Instinct GPUs
1 source · show source ▾
TILERT DECODE · Jul 14
vLLM adds a TileRT-powered latency-optimized decode engine
1 source · show source ▾
MOE DISAGGREGATION · Jul 23
vLLM ships an AFD plugin, disaggregating attention and FFN for MoE serving
1 source · show source ▾
NOW · Jul 27
Netflix details its in-house Triton + vLLM serving platform
1 source · show source ▾

What to watch — open questions

  • Does Netflix's in-house platform use any of the three July engine upgrades (EAGLE-3, TileRT, AFD), or run a separate internal stack on top of vLLM?
  • Do the EAGLE-3 AMD throughput gains and AFD's Ascend NPU backend hold up on independent benchmarks outside vLLM's own posts?
How this thread was built
editor wrote the arc · 4 beats

Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.