Story
vllm_releases · Jul 11, 2026 · release
Source brief
vllm v0.25.0
github.comJul 11, 2026
original source linked
Release highlights
- Model Runner V2 is now the default for all dense models . Building on quantized-model support from the previous release, MRv2 is now the standard execution p...
- PagedAttention has been removed . The legacy attention implementation is deleted now that V1/MRv2 backends are the standard path
- The Transformers modeling backend is now as fast as native vLLM , and gained FP8 MoE support , CUDA graph + embed scaling fixes , and migration of GPTBigCode...
Feed lens
agentic