LLM Digest
Subscribe

AI Storyline

3 items · 2 sources · 3 days

View as JSON

Operational story trace

Vllm Serving

Latest change

Netflix disclosed the internals of its in-house LLM serving platform, built on Triton and vLLM — the thread's first large-scale production-adoption proof point after vLLM's July engine upgrades.

Earlier contextThe story so far

vLLM shipped a TileRT-powered latency-optimized decode engine that pairs with existing prefill behind one shared serving layer, then an Attention-FFN Disaggregation plugin adding GPU and Ascend NPU backends for MoE serving. Both landed as drop-in additions through the same V1 connector interface, with no changes required to already-running deployments.

editor-curated · source-linked

Arc

Jul 14Jul 27 · now
TILERT DECODE · Jul 14
vLLM adds a TileRT-powered latency-optimized decode engine
1 source · show source ▾
MOE DISAGGREGATION · Jul 23
vLLM ships an AFD plugin, disaggregating attention and FFN for MoE serving
1 source · show source ▾
NOW · Jul 27
Netflix details its in-house Triton + vLLM serving platform
1 source · show source ▾

What to watch — open questions

  • Does Netflix's in-house platform build on vLLM's TileRT decode and AFD disaggregation upgrades, or run a separate internal stack on top of vLLM?
  • Does AFD's Ascend NPU backend and TileRT's decode latency gains hold up on independent benchmarks outside vLLM's own posts?
How this thread was built
editor wrote the arc · 3 beats

Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.