Story
vllm_blog · Jul 22, 2026 · news
vllm.aiJul 22, 2026
original source linked
In brief
A preview of production-scale Kimi K3 support in vLLM, including KDA-aware prefix caching, fused kernels, optimized MXFP4 MoE, multimodal integration, and initial NVIDIA and AMD paths.
Continue reading