RepoJournal
Local LLMs Local LLMs
69 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-09-22
stories 275

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

vLLM 0.30.0 lands DeepSeek-V4.1-Flash, and a min_tokens validation hole closes

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

vLLM's 762-commit release and its min_tokens rejection headline a day where llama.cpp tuned every accelerator backend except the default one.

v0.30.0 vllm-project/vllm

by khluu

762 commits from 315 contributors, with DeepSeek-V4.1-Flash storing its whole KV cache in MXFP8 through the FlashMLA V4.1 record on SM100, plus DeepSeek-V4-Flash-Vision-Exp, GLM-5.3-Flash with EPLB, and K2-Horizon.

[Security] Reject min_tokens that exceeds the filled max_tokens default (#57731) vllm-project/vllm

by Juan Pérez de Algaba

Requests whose min_tokens exceeds the filled default max_tokens are now rejected rather than accepted, per the commit from Juan Pérez de Algaba. Treat this as behavior your clients must not depend on: pin the version and audit callers that set min_tokens without setting max_tokens.

hexagon: overhaul of buffer and DMA handling to support 64bit mappings and general improvements ggml-org/llama.cpp

by max-krasnyansky

Hexagon v81 and up (Gen5, X2-Elite, IQ10) can now map buffers above the 4GB virtual address space on the NPU, usable only with the DMA engine, which the author says lifts performance on larger models like gemma-4-26B and gpt-oss-20b.

[Bugfix][GDN] Fix stateless first-chunk classification vllm-project/vllm

by taking-lying-flat

GDNMetadataBuilder.build() passed decode_threshold=1 while leaving treat_short_extends_as_decodes=True, so every one-token sequence was classified as a decode, including the first chunk of a request with no prior GDN state; since Mamba-style state pages can be reused unzeroed, that first token could read stale state. The decode path now classifies the stateless first chunk correctly.

ggml-metal : simplify fusion pattern op list declaration ggml-org/llama.cpp

by ggerganov

The Metal fusion table in ggml-metal-fusion.cpp now declares only the full raw op sequence, with the non-empty sequence derived once at static initialization; duplicate top-k MoE and MoE reduce lists are gone, with no runtime behavior change.

[AMD] [GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform sgl-project/sglang

by Jacob0226

GLM-5.3-Flash hands a top-k width of 2051 (512 pools expanded to 2048 plus up to 3 leftovers) to a DSA page-table transform that asserted exactly 2048, a path the default single combined kernel never reaches but three supported split setups do. Quiet day elsewhere: a ROCm cast removal in the jit grouped topk path, a TensorCast HiCache backend, a use-scoped layerwise release split in the diffusi...

Quick answers

What shipped in Local LLMs on September 22, 2026?
vLLM's 762-commit release and its min_tokens rejection headline a day where llama.cpp tuned every accelerator backend except the default one. In total, 132 commits, 132 pull requests, and 11 releases landed.
Who contributed to Local LLMs on September 22, 2026?
14 developers shipped this update, including yomaytk, max-krasnyansky, pl752, ggerganov, khluu, Juan Pérez de Algaba, BugenZhao, and taking-lying-flat, and 6 more.
What were the notable Local LLMs updates?
v0.30.0, [Security] Reject min_tokens that exceeds the filled max_tokens default (#57731), and hexagon: overhaul of buffer and DMA handling to support 64bit mappings and general improvements.