vLLM x TileRT: Specialized Decode for Latency-Critical Serving
vLLM pairs prefill with a separate TileRT decode engine behind one shared serving layer, with no changes needed to existing deployments.
3 items · 2 sources · 3 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
Netflix disclosed the internals of its in-house LLM serving platform, built on Triton and vLLM — the thread's first large-scale production-adoption proof point after vLLM's July engine upgrades.
vLLM shipped a TileRT-powered latency-optimized decode engine that pairs with existing prefill behind one shared serving layer, then an Attention-FFN Disaggregation plugin adding GPU and Ascend NPU backends for MoE serving. Both landed as drop-in additions through the same V1 connector interface, with no changes required to already-running deployments.
Arc
vLLM pairs prefill with a separate TileRT decode engine behind one shared serving layer, with no changes needed to existing deployments.
vLLM's AFD plugin disaggregates attention and FFN computation for MoE models, adding GPU and Ascend NPU backend support.
Netflix goes public with its in-house LLM serving platform built on Triton and vLLM — the thread's first large-scale production-adoption signal.
vLLM pairs prefill with a separate TileRT decode engine behind one shared serving layer, with no changes needed to existing deployments.
vLLM's AFD plugin disaggregates attention and FFN computation for MoE models, adding GPU and Ascend NPU backend support.
Netflix goes public with its in-house LLM serving platform built on Triton and vLLM — the thread's first large-scale production-adoption signal.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.