LLM Digest
Subscribe

Story

vllm_releases · Jul 27, 2026 · release

Source brief

vllm v0.26.0

github.comJul 27, 2026
original source linked

Release highlights

  • New Inkling model family with a full support stack: base modeling , piecewise CUDA graph support , Hopper FA4 relative attention , MTP=1 speculative decoding...
  • DeepSeek-V4 performance push across vendors: a specialized routing kernel (2.94% E2E TPOT, #48660 ), fused_topk_bias (1.5–2x kernel, #47463 ), and redundant...
  • fp32 lm_head for generation models via head_dtype , extended to the LoRA path and given a ROCm torch.mm fast path , improving accuracy for generation heads
Feed lens
eval

Continue reading

Read the original at github.com →Open in live feedRead that day’s brief

Earlier in this thread 4 items