$ the-wire · showcase
Qwen4Exp MTP lands in llama.cpp as SGLang ships v0.5.21
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
llama.cpp added multi-token prediction for Qwen4Exp and repaired its attention path, SGLang pushed a 779-PR release with new model support, and vLLM's ROCm path got both a prefill speedup and a revert that unbreaks MiniMax-M3.
Qwen4Exp: add MTP ggml-org/llama.cpp
MTP adds speculative decoding to Qwen4Exp; on a small speed-bench slice on DGX Spark with Qwen3.8-Flash-Next iq4_xs, coding decode went from 28.27 to 41.91 t/s under --spec-type draft-mtp with --spec-draft-n-max 3. Server operators can expect higher decode throughput on that model family once built.
llama: fix qwen4exp ggml-org/llama.cpp
Qwen4exp was broken on current master, and this fixes the attention path around how llama-memory-hybrid-idx pooling works. If you are pulling master for Qwen3.8-Flash-Next, take this build rather than the one before it.
v0.5.21 sgl-project/sglang
779 PRs from 227 contributors in one release, adding DeepSeek-V4.1 Flash, GigaChat 3.5, IQuest-Q1 and MiMo-V2.6/MiMo-V2.6-Pro across LLM and VLM modes. Treat it as a breaking release and read the notes before upgrading a serving cluster.
[Bugfix][Frontend] Accept Anthropic tool_addition and tool_removal content blocks vllm-project/vllm
Claude Code sends tool_addition and tool_removal blocks inside role: "system" messages when tool search is on, and /v1/messages answered 400, so the client retried without them. Resolving those blocks into the tool list the chat template sees matters because defer_loading: true tools stay hidden in deferral-aware templates such as GLM-5.1.
[ROCm][BugFix] Revert AITER PA gluon decode from ROCM_AITER_FA vllm-project/vllm
ROCm's AITER gluon decode kernel needs a uniform query length, but the scheduler emits ragged decode batches whenever a speculative draft is truncated or a token budget clips a request; MiniMax-M3 with EAGLE3 and shuffle KV layout aborted on the first ragged step. The pa gluon changes in rocm_aiter_fa.py are reverted, with a follow-up promised to restore them.
models: add clef support via llama-server ollama/ollama
Ollama added clef model support routed through llama-server, and a separate change reports only the "decision" capability for models that declare it, so clients stop offering them for general chat, tools, or thinking while scheduling and serving still check the full capability set. MLX also landed four fixes: reductions over views past 2^31, bare-ellipsis __setitem__, fft size one, and a MultiO...
Action items