Story
vllm_releases · Jun 30, 2026 · release
Source brief
vllm v0.24.0
github.comJun 30, 2026
original source linked
Release highlights
- MiniMax-M3 : Added support for the new MiniMax-M3 model , with a fast follow-on of BF16/FP8 indexer via MSA , MXFP4 support , FP8 sparse GQA , and extensive...
- DeepSeek-V4 keeps maturing : Following its debut, DeepSeek-V4 received another large optimization pass — a FlashInfer sparse index cache (2–4% TTFT) , prefil...
- Model Runner V2 (MRv2) continues to expand : MRv2 now supports quantized models by default , enables GraniteMoE by default , and gained migration of Qwen + D...
Feed lens
eval