vLLM v0.28.0: Major Performance Push for Kimi-K3 and DeepSeek V4
vLLM v0.28.0 ships with 584 commits from 270 contributors, focusing on major performance optimizations for Kimi-K3 and DeepSeek V4. Key highlights include Decode Context Parallel (DCP), fused FlashKDA kernels, SiTU activation for MegaMoE, and adaptive speculative token budgets, delivering up to 60% better TTFT. DeepSeek V4 gains sparse MLA end-to-end support, AMD Quark NVFP4, and ROCm enablement. The release also matures Model Runner V2 with E/P/D disaggregation, introduces tiered KV cache offloading to disk, and adds a Rust frontend with gRPC multimodal inference. New defaults raise max_num_batched_tokens to 16384, and several breaking changes are noted, including bitsandbytes moving to a plugin.
Kimi-K3 also now runs on ROCm with the V2 model runner.
- joshheitzman
I was hoping to see the reasoning_content mess get robustly fixed, but all we got was this doc change: https://github.com/vllm-project/vllm/pull/50624
- SillyUsername
I just wish they'd support Pascal :(
Nvidia might have given up support but it doesn't mean vllm have to (llama.cpp didn't).
- kouteiheika
I love vLLM, but damn if it isn't frustratingly buggy.
I was recently running DeepSeek-V4-Flash on a B300. On v0.26 it was totally broken, and I had to add three out-of-tree patches to fix it. I updated to v0.27 -- no patches necessary now, but the output is now broken as it randomly starts responding with garbage (repeated token loops). On my workstation where I run Gemma-4 on an RTX 6000 the whole process tends to get stuck and stops responding, and needs to be killed and restarted to start working again. On my friend's 4x RTX 6000 box where he runs DeepSeek-V4-Flash high concurrency also triggers some kind of a bug where it spews out garbage, but this time it's not a single repeated token and looks like this: (this is copy-pasted from what the model did output, genuinely looks like it was in pain trying to end its thinking trace but not being able to)
<|beginofsentence|>| only text. No. I<|beginofsentence|>#done. Whatever. Do<|beginofsentence|>### final response.Content-EncodingDone.</ /div> It's over? Let's this.No, Em,Okay, finally.<|beginofsentence|>import re and I can<|beginofsentence|>È.Let's finish? Next.Content No matter)2Stay.2No. Childish.No No commentsrandom DoBaBye-. ... Whatever. Alright.</body>No And 2. about: So)</think> No text I'm tired This is fine)))ExodusNo comments, hiddenNo. Nonsense Mehski. N-Hay que noIbye. Goodbye Last sentence after(ok copy Nothing useful. OK OK. . . . .Come on Nancy . . . . . . Let's just end this please.No matter what. about […]
- zoobab
Did some loadtests on vllm, managed to crash it :-)
- Der_Einzige
Still way behind on LLM sampler support compared to llama-cpp. Where's support for top-n-sigma? for DRY? for XTC? C'mon guys!