$ the-wire · showcase
vLLM 0.30.0 lands DeepSeek-V4.1-Flash, and a min_tokens validation hole closes
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
vLLM's 762-commit release and its min_tokens rejection headline a day where llama.cpp tuned every accelerator backend except the default one.
v0.30.0 vllm-project/vllm
762 commits from 315 contributors, with DeepSeek-V4.1-Flash storing its whole KV cache in MXFP8 through the FlashMLA V4.1 record on SM100, plus DeepSeek-V4-Flash-Vision-Exp, GLM-5.3-Flash with EPLB, and K2-Horizon.
[Security] Reject min_tokens that exceeds the filled max_tokens default (#57731) vllm-project/vllm
Requests whose min_tokens exceeds the filled default max_tokens are now rejected rather than accepted, per the commit from Juan Pérez de Algaba. Treat this as behavior your clients must not depend on: pin the version and audit callers that set min_tokens without setting max_tokens.
hexagon: overhaul of buffer and DMA handling to support 64bit mappings and general improvements ggml-org/llama.cpp
Hexagon v81 and up (Gen5, X2-Elite, IQ10) can now map buffers above the 4GB virtual address space on the NPU, usable only with the DMA engine, which the author says lifts performance on larger models like gemma-4-26B and gpt-oss-20b.
[Bugfix][GDN] Fix stateless first-chunk classification vllm-project/vllm
GDNMetadataBuilder.build() passed decode_threshold=1 while leaving treat_short_extends_as_decodes=True, so every one-token sequence was classified as a decode, including the first chunk of a request with no prior GDN state; since Mamba-style state pages can be reused unzeroed, that first token could read stale state. The decode path now classifies the stateless first chunk correctly.
ggml-metal : simplify fusion pattern op list declaration ggml-org/llama.cpp
The Metal fusion table in ggml-metal-fusion.cpp now declares only the full raw op sequence, with the non-empty sequence derived once at static initialization; duplicate top-k MoE and MoE reduce lists are gone, with no runtime behavior change.
[AMD] [GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform sgl-project/sglang
GLM-5.3-Flash hands a top-k width of 2051 (512 pools expanded to 2048 plus up to 3 leftovers) to a DSA page-table transform that asserted exactly 2048, a path the default single combined kernel never reaches but three supported split setups do. Quiet day elsewhere: a ROCm cast removal in the jit grouped topk path, a TensorCast HiCache backend, a use-scoped layerwise release split in the diffusi...
Action items
- → Upgrade vllm to 0.30.0 to pick up the min_tokens rejection, and audit clients that set min_tokens without an explicit... vllm-project/vllm [immediate]