{"slug":"serving-vllm","label":"Vllm Serving","item_count":3,"day_count":3,"source_count":2,"first_seen":"2026-07-14T00:00:00+00:00","last_updated":"2026-07-27T07:33:13+00:00","generated_at":"2026-08-03T20:06:24.210843+00:00","sources":["search_llm_ops_news","vllm_blog"],"days":[{"date":"2026-07-14","items":[{"title":"vLLM x TileRT: Specialized Decode for Latency-Critical Serving","url":"https://vllm.ai/blog/2026-07-14-vllm-tilert-pd","source":"vllm_blog","type":"news","summary_1line":"vLLM prefill paired with TileRT decode through vLLM V1's connector interface: a specialized, latency-optimized decode engine that coexists with native vLLM decode behind one shared serving layer, with zero changes to...","sid":"d08095949d6300c2","published":"2026-07-14T00:00:00+00:00","editor_note":"vLLM pairs prefill with a separate TileRT decode engine behind one shared serving layer, with no changes needed to existing deployments."}]},{"date":"2026-07-23","items":[{"title":"Announcing vLLM AFD Plugin: Disaggregating Attention and FFN for Flexible MoE Serving","url":"https://vllm.ai/blog/2026-07-23-vllm-afd-plugin","source":"vllm_blog","type":"news","summary_1line":"What vLLM AFD Plugin adds to the vLLM ecosystem: Attention–FFN disaggregation for MoE serving, GPU and Ascend NPU backends, connector-based execution, and graph and ubatching support.","sid":"94813f8b6bc86093","published":"2026-07-23T00:00:00+00:00","editor_note":"vLLM's AFD plugin disaggregates attention and FFN computation for MoE models, adding GPU and Ascend NPU backend support."}]},{"date":"2026-07-27","items":[{"title":"Netflix Details Its In-House LLM Serving Platform with Triton and vLLM - infoq.com","url":"https://news.google.com/rss/articles/CBMiZ0FVX3lxTFBpby0wOTVaSTc0LWVRam9WelpQRDdhdW1KMGJYWFRIRWZnQXFGYWdhVjAxUGRhNkRlRGl4VEZnRFVnbEVqeWJtdGVqSFlNa1BIcVhtZzBqdFZHTHlQS3daQXdJRmdrQ0U?oc=5","source":"search_llm_ops_news","type":"news","summary_1line":"Netflix Details Its In-House LLM Serving Platform with Triton and vLLM infoq.com","sid":"64c163bb191bab4e","published":"2026-07-27T07:33:13+00:00","editor_note":"Netflix goes public with its in-house LLM serving platform built on Triton and vLLM — the thread's first large-scale production-adoption signal."}]}],"editorial":{"tldr":"vLLM shipped a TileRT-powered latency-optimized decode engine that pairs with existing prefill behind one shared serving layer, then an Attention-FFN Disaggregation plugin adding GPU and Ascend NPU backends for MoE serving. Both landed as drop-in additions through the same V1 connector interface, with no changes required to already-running deployments.","stale":false,"whats_new":"Netflix disclosed the internals of its in-house LLM serving platform, built on Triton and vLLM — the thread's first large-scale production-adoption proof point after vLLM's July engine upgrades.","why_it_matters":"These are incremental, drop-in performance levers rather than rewrites — teams already running vLLM can adopt TileRT decode or AFD disaggregation one at a time, and Netflix's disclosure shows the connector model already scales to hyperscaler traffic.","take_for_builders":"If you're serving MoE or long-context workloads on vLLM, benchmark the AFD plugin and TileRT decode path against your current setup before your next capacity review — both are additive through the V1 connector, so you can pilot without a stack rewrite.","beats":[{"kicker":"TILERT DECODE","tone":"launch","headline":"vLLM adds a TileRT-powered latency-optimized decode engine","summary":"vLLM pairs its existing prefill with a separate TileRT decode engine through the V1 connector interface, coexisting with native decode behind one shared serving layer.","sids":["d08095949d6300c2"]},{"kicker":"MOE DISAGGREGATION","tone":"rising","headline":"vLLM ships an AFD plugin, disaggregating attention and FFN for MoE serving","summary":"The Attention-FFN Disaggregation plugin adds GPU and Ascend NPU backend support with connector-based execution, graph mode, and ubatching for MoE models.","sids":["94813f8b6bc86093"]},{"kicker":"NOW","tone":"now","headline":"Netflix details its in-house Triton + vLLM serving platform","summary":"Netflix goes public with how its production LLM serving stack combines Triton and vLLM — the thread's first big real-world adoption signal, not just an engine feature.","sids":["64c163bb191bab4e"]}],"open_questions":["Does Netflix's in-house platform build on vLLM's TileRT decode and AFD disaggregation upgrades, or run a separate internal stack on top of vLLM?","Does AFD's Ascend NPU backend and TileRT's decode latency gains hold up on independent benchmarks outside vLLM's own posts?"],"generated_at":"2026-08-03T00:04:27+00:00"}}