RepoJournal
Local LLMs Local LLMs
71 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-10-02
stories 281

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

Qwen4Exp MTP lands in llama.cpp as SGLang ships v0.5.21

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

llama.cpp added multi-token prediction for Qwen4Exp and repaired its attention path, SGLang pushed a 779-PR release with new model support, and vLLM's ROCm path got both a prefill speedup and a revert that unbreaks MiniMax-M3.

Qwen4Exp: add MTP ggml-org/llama.cpp

by am17an

MTP adds speculative decoding to Qwen4Exp; on a small speed-bench slice on DGX Spark with Qwen3.8-Flash-Next iq4_xs, coding decode went from 28.27 to 41.91 t/s under --spec-type draft-mtp with --spec-draft-n-max 3. Server operators can expect higher decode throughput on that model family once built.

llama: fix qwen4exp ggml-org/llama.cpp

by am17an

Qwen4exp was broken on current master, and this fixes the attention path around how llama-memory-hybrid-idx pooling works. If you are pulling master for Qwen3.8-Flash-Next, take this build rather than the one before it.

v0.5.21 sgl-project/sglang

by Fridge003

779 PRs from 227 contributors in one release, adding DeepSeek-V4.1 Flash, GigaChat 3.5, IQuest-Q1 and MiMo-V2.6/MiMo-V2.6-Pro across LLM and VLM modes. Treat it as a breaking release and read the notes before upgrading a serving cluster.

[Bugfix][Frontend] Accept Anthropic tool_addition and tool_removal content blocks vllm-project/vllm

by Anurag-M1

Claude Code sends tool_addition and tool_removal blocks inside role: "system" messages when tool search is on, and /v1/messages answered 400, so the client retried without them. Resolving those blocks into the tool list the chat template sees matters because defer_loading: true tools stay hidden in deferral-aware templates such as GLM-5.1.

[ROCm][BugFix] Revert AITER PA gluon decode from ROCM_AITER_FA vllm-project/vllm

by ukannika

ROCm's AITER gluon decode kernel needs a uniform query length, but the scheduler emits ragged decode batches whenever a speculative draft is truncated or a token budget clips a request; MiniMax-M3 with EAGLE3 and shuffle KV layout aborted on the first ragged step. The pa gluon changes in rocm_aiter_fa.py are reverted, with a follow-up promised to restore them.

models: add clef support via llama-server ollama/ollama

by jmorganca

Ollama added clef model support routed through llama-server, and a separate change reports only the "decision" capability for models that declare it, so clients stop offering them for general chat, tools, or thinking while scheduling and serving still check the full capability set. MLX also landed four fixes: reductions over views past 2^31, bare-ellipsis __setitem__, fft size one, and a MultiO...

Quick answers

What shipped in Local LLMs on October 2, 2026?
llama.cpp added multi-token prediction for Qwen4Exp and repaired its attention path, SGLang pushed a 779-PR release with new model support, and vLLM's ROCm path got both a prefill speedup and a revert that unbreaks MiniMax-M3. In total, 135 commits, 135 pull requests, and 11 releases landed.
Who contributed to Local LLMs on October 2, 2026?
17 developers shipped this update, including jmorganca, dhiltgen, Aman Gupta, njsyw1997, kliuae, ukannika, Anurag-M1, and sfeng33, and 9 more.
What were the notable Local LLMs updates?
Qwen4Exp: add MTP, llama: fix qwen4exp, and v0.5.21.