LLM Digest
Subscribe

Story

vllm_releases · Jun 30, 2026 · release

Source brief

vllm v0.24.0

github.comJun 30, 2026
original source linked

Release highlights

  • MiniMax-M3 : Added support for the new MiniMax-M3 model , with a fast follow-on of BF16/FP8 indexer via MSA , MXFP4 support , FP8 sparse GQA , and extensive...
  • DeepSeek-V4 keeps maturing : Following its debut, DeepSeek-V4 received another large optimization pass — a FlashInfer sparse index cache (2–4% TTFT) , prefil...
  • Model Runner V2 (MRv2) continues to expand : MRv2 now supports quantized models by default , enables GraniteMoE by default , and gained migration of Qwen + D...
Feed lens
eval

Continue reading

Read the original at github.com →Open in live feedRead that day’s brief

Earlier in this thread 4 items