LLM Digest
Subscribe

Story

vllm_blog · Sep 21, 2026 · news

Source brief

PD Serving of Qwen3.8-2.4T

vllm.aiSep 21, 2026
original source linked

In brief

How vLLM reaches 5K throughput and 180 interactivity on Qwen3.8-2.4T with GB300 NVL72 PD serving and how to reproduce results yourself.

Continue reading

Read the original at vllm.ai →Open in live feed

Earlier in this thread 4 items