RepoJournal
Local LLMs Local LLMs
72 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-10-03
stories 262

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

llama.cpp spec sampling goes probabilistic, vLLM trims GLM-5.3 attention overhead

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

llama.cpp reworks speculative decoding to stop discarding the drafter's distribution, while vLLM and sglang land targeted kernel and cache fixes that change what existing deployments do at temperature above zero and after prefix eviction.

Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP ggml-org/llama.cpp

by praneshgo

Model-based drafters used to run the sampler, throw the result away, and hand the target a single top candidate, which the target then had to draw again at temperature above zero; the drafter is now probabilistic and the target verifies by rejection sampling, with a flag to enable probabilistic draft sampling (default stays greedy) and argmax fallback for grammar-constrained requests. The PR st...

[GLM5.3 Perf] Reuse sparse MLA index conversion across layers, 3.5~3.9x kernel performance improvement vllm-project/vllm

by yewentao256

GLM-5.3 shares top-k indices across layer groups, but FA3 sparse MLA was re-deriving the same physical KV slots per layer; converting once per group cuts conversions from 78 to 21 per main-model forward, cited as a 3.5 to 3.9x kernel improvement.

fix: preserve QSA indexer state through HiCache sgl-project/sglang

by alphabetc1

QSA compressed index keys live outside full_kv_pool, so the hybrid HiCache stack was restoring KV and Mamba state without the indexer state that selects attention blocks. After device eviction that let a prefix hit read valid KV through stale QSA block selections and produce incorrect output, a correctness fix for anyone running prefix caching on the hybrid stack. The related PLE slot-state fix...

CUDA: fuse shared experts into MMVQ ggml-org/llama.cpp

by am17an

Shared experts are now fused into the MMVQ path; on Qwen3.5-35B-A3B-Q8_0 on RTX Pro 6000 the gain runs from +0.42% at microbatch 4 to +5.06% at microbatch 5, with +4.54% at microbatch 2.

Size the VMM graph-input exchange by the widest input across ranks sgl-project/sglang

by metamergebot

The VMM graph-input exchange packed expandable-segment chunk indices into a fixed-width struct sized for 16 chunks, so a rank whose input spanned more raised "Too many VMM chunks for graph input" while its peers blocked in the all-gather. The exchange is now sized by the widest input across ranks, which removes a hang that depended on allocator segment size rather than the model.

Quick answers

What shipped in Local LLMs on October 3, 2026?
llama.cpp reworks speculative decoding to stop discarding the drafter's distribution, while vLLM and sglang land targeted kernel and cache fixes that change what existing deployments do at temperature above zero and after prefix eviction. In total, 126 commits, 126 pull requests, and 10 releases landed.
Who contributed to Local LLMs on October 3, 2026?
16 developers shipped this update, including ngxson, Pranesh Gonegandla, am17an, Dovis01, yewentao256, stecasta, mgoin, and kkHuang-amd, and 8 more.
What were the notable Local LLMs updates?
Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP, [GLM5.3 Perf] Reuse sparse MLA index conversion across layers, 3.5~3.9x kernel performance improvement, and fix: preserve QSA indexer state through HiCache.