RepoJournal
Local LLMs Local LLMs
76 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-10-05
stories 193

© 2026 RepoJournal Home Showcase How it works Privacy

$ the-wire · showcase

vLLM 0.31.0 ships DeepSeek-V4.1-Flash kernels, streaming errors stop vanishing behind DONE

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

vLLM's 0.31.0 lands 717 commits of DeepSeek-V4.1-Flash perf work while llama.cpp, sglang, and MLX each close correctness gaps that were silently costing you tokens or throughput.

v0.31.0 vllm-project/vllm

by khluu

717 commits from 307 contributors, 96 of them new, with DeepSeek-V4.1-Flash at the center: FlashMLA mega attention on the V4.1 NVFP4 compressed KV cache is now the SM100 default, joined by DeepGEMM sparse MQA logits for the indexer and a Mega-Gate that fuses the gate GEMM with expert selection. Decoder boundaries now fuse the TP all-reduce, mHC input preparation, and the MoE finalize, so SM100 ...

[Bugfix][Frontend] Return streaming errors before the first token vllm-project/vllm

by 1fanwang

When chat or completion generation failed before emitting a token, the empty-chunk filter threw away the terminal error and clients got a bare DONE with no error at all. The generators now check for errors before filtering empty chunks, so both APIs return an SSE 500 before DONE. This fixes the streaming half of issue 40978; 41929 adds a KV-load stop reason but still checks the error after the ...

llama: support both embd + raw tokens in batch ggml-org/llama.cpp

by ngxson

llama_batch_ext now accepts both embedding vectors and text tokens in the same batch, which is what non-causal prompts like paligemma need. Models that can't handle it are gated internally by llm_arch_supports_mixed_batch and return an error instead of misbehaving; mtmd handling is still outstanding.

Use the fast qmv kernel for outputs not divisible by 8 ml-explore/mlx

by wyanzhao

QMV used to fall back to the generic kernel whenever the output size wasn't divisible by 8, even with an input that met the qmv_fast alignment requirement, which is exactly the shape of the narrow N=1 and N=4 projections in autoregressive decode. New affine_qmv_fast_rows and fp_qmv_fast_rows kernels pad the final SIMD group with the last valid weight row and store only valid rows.

metal : few-row MMA mat-mul and batched copies for speculative decoding ggml-org/llama.cpp

by pratiknarola-t

On M1 through M4 there's no tensor API, so speculative decoding's 2 to 16 src1 rows ran through mat-vec kernels whose cost grows with each row; on an M3 Ultra that made DFlash2 decoding of Qwen3.8-27B slower than serial decoding on master. This adds mat-mul kernels for 2 to 16 src1 rows on 8x8 simdgroup matrices plus batched copies.

Fix ResidencySets reading autoreleased error object (#4623) ml-explore/mlx

by LongYinan

Elsewhere: MLX fixed ResidencySets reading an autoreleased error object and a CPU float16 NEON comparison bug that broadcast the first lane's mask across mixed-value vectors, adding JVP support for logcumsumexp, llama.cpp rewrote its logger for self-contained colors and passed router color settings to children, and sglang consolidated four copies of the in-flight prefetch teardown into one helper.

Quick answers

What shipped in Local LLMs on October 5, 2026?
vLLM's 0.31.0 lands 717 commits of DeepSeek-V4.1-Flash perf work while llama.cpp, sglang, and MLX each close correctness gaps that were silently costing you tokens or throughput. In total, 91 commits, 91 pull requests, and 11 releases landed.
Who contributed to Local LLMs on October 5, 2026?
20 developers shipped this update, including ngxson, ServeurpersoCom, JohannesGaessler, SongXiaoXi, pratiknarola-t, khluu, yongqiw-i, and 1fanwang, and 12 more.
What were the notable Local LLMs updates?
v0.31.0, [Bugfix][Frontend] Return streaming errors before the first token, and llama: support both embd + raw tokens in batch.