LLM Digest
Subscribe

Story

vllm_releases · Aug 26, 2026 · release

Source brief

vllm v0.28.0

github.comAug 26, 2026
original source linked

Release highlights

  • Kimi-K3 performance push : a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support , fused FlashKDA decode and prefi...
  • DeepSeek V4 : sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding , joined by AMD Quark NVFP4 support , reasoning-effort p...
  • Speculative decoding advances : DFlash2 with local convolution and a candidate selector , DSpark confidence-scheduled verification , and async scheduling aut...
Feed lens
codex

Continue reading

Read the original at github.com →Open in live feedRead that day’s brief

Earlier in this thread 4 items