Story
vllm_releases · Aug 26, 2026 · release
Source brief
vllm v0.28.0
github.comAug 26, 2026
original source linked
Release highlights
- Kimi-K3 performance push : a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support , fused FlashKDA decode and prefi...
- DeepSeek V4 : sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding , joined by AMD Quark NVFP4 support , reasoning-effort p...
- Speculative decoding advances : DFlash2 with local convolution and a candidate selector , DSpark confidence-scheduled verification , and async scheduling aut...
Feed lens
codex