$ the-wire · showcase
vLLM 0.31.0 ships DeepSeek-V4.1-Flash kernels, streaming errors stop vanishing behind DONE
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
vLLM's 0.31.0 lands 717 commits of DeepSeek-V4.1-Flash perf work while llama.cpp, sglang, and MLX each close correctness gaps that were silently costing you tokens or throughput.
v0.31.0 vllm-project/vllm
717 commits from 307 contributors, 96 of them new, with DeepSeek-V4.1-Flash at the center: FlashMLA mega attention on the V4.1 NVFP4 compressed KV cache is now the SM100 default, joined by DeepGEMM sparse MQA logits for the indexer and a Mega-Gate that fuses the gate GEMM with expert selection. Decoder boundaries now fuse the TP all-reduce, mHC input preparation, and the MoE finalize, so SM100 ...
[Bugfix][Frontend] Return streaming errors before the first token vllm-project/vllm
When chat or completion generation failed before emitting a token, the empty-chunk filter threw away the terminal error and clients got a bare DONE with no error at all. The generators now check for errors before filtering empty chunks, so both APIs return an SSE 500 before DONE. This fixes the streaming half of issue 40978; 41929 adds a KV-load stop reason but still checks the error after the ...
llama: support both embd + raw tokens in batch ggml-org/llama.cpp
llama_batch_ext now accepts both embedding vectors and text tokens in the same batch, which is what non-causal prompts like paligemma need. Models that can't handle it are gated internally by llm_arch_supports_mixed_batch and return an error instead of misbehaving; mtmd handling is still outstanding.
Use the fast qmv kernel for outputs not divisible by 8 ml-explore/mlx
QMV used to fall back to the generic kernel whenever the output size wasn't divisible by 8, even with an input that met the qmv_fast alignment requirement, which is exactly the shape of the narrow N=1 and N=4 projections in autoregressive decode. New affine_qmv_fast_rows and fp_qmv_fast_rows kernels pad the final SIMD group with the last valid weight row and store only valid rows.
metal : few-row MMA mat-mul and batched copies for speculative decoding ggml-org/llama.cpp
On M1 through M4 there's no tensor API, so speculative decoding's 2 to 16 src1 rows ran through mat-vec kernels whose cost grows with each row; on an M3 Ultra that made DFlash2 decoding of Qwen3.8-27B slower than serial decoding on master. This adds mat-mul kernels for 2 to 16 src1 rows on 8x8 simdgroup matrices plus batched copies.
Fix ResidencySets reading autoreleased error object (#4623) ml-explore/mlx
Elsewhere: MLX fixed ResidencySets reading an autoreleased error object and a CPU float16 NEON comparison bug that broadcast the first lane's mask across mixed-value vectors, adding JVP support for logcumsumexp, llama.cpp rewrote its logger for self-contained colors and passed router color settings to children, and sglang consolidated four copies of the in-flight prefetch teardown into one helper.