RepoJournal
Local LLMs Local LLMs
69 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-09-21
stories 184

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

llama.cpp sanitizes invalid UTF-8 in the parser AST

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

Parser robustness, CUDA attention tuning, and memory bounds dominate a day of work aimed at the long tail of production inference: malformed model output, sparse attention on Qwen4, and an indexer that OOMs with room to spare in the cache.

common/peg : handle invalid utf-8 sequences in the AST ggml-org/llama.cpp

by aldehir

Models occasionally emit invalid UTF-8, which used to fail during parsing; the parser now records those runs and exposes sanitized_text() on every AST node, replacing invalid sequences with the maximal subpart per Unicode recommendations. Tool names and args are deliberately left un-sanitized on the assumption that grammar constraints keep them valid, so anything feeding tool call text through ...

CUDA: enable sparse fa for qwen4 ggml-org/llama.cpp

by am17an

Sparse FlashAttention is now enabled for Qwen4 in the CUDA path, switching on once context reaches 32768 by taking the union of tokens in use at ncols1=8. The author notes the model's attention is otherwise unoptimized, with the entire kv-cache re-scored on every step, and flags that as still open.

[DeepSeek-V4.1] Bound dense prefill indexer memory sgl-project/sglang

by harmya

The DeepSeek-V4.1 dense prefill indexer could OOM even with KV cache headroom, because DeepGEMM returns FP32 scores shaped [query tokens, visible context entries] that stay alive while Boolean candidate masks are built in pieces and concatenated. The fix bounds that memory; FP4 inputs are unaffected in precision.

Use pinned memory for asynchronous sampling metadata transfers sgl-project/sglang

by zRzRzRzRzRzRzR

Sampling metadata now moves over pinned host memory: stage penalty parameters, sparse logit bias, custom-processor masks, and raw token-ID logprob indices use nonblocking device copies, and stop-token IDs are padded on CPU and transferred once per batch. Sampling semantics, dtypes, device placement, and platform-specific pinning support are preserved, so this is a throughput change rather than ...

metal : support arbitrary hc in dsv4_hc_pre ggml-org/llama.cpp

by ggerganov

Metal's dsv4_hc_pre kernels now accept arbitrary hc values instead of the hardcoded hc = 4. Elsewhere in the day's long tail: CUDA FlashAttention tuning for Gemma 4 head sizes 256 and 512, CUDA graph capture for the DeepSeek-V4.1 vision encoder, Humming quantization gaining Hadamard transforms and SM120/SM121 support, a dead-kernel cleanup, an NVFP4 MoE dispatch fix where create_weights and app...

Quick answers

What shipped in Local LLMs on September 21, 2026?
Parser robustness, CUDA attention tuning, and memory bounds dominate a day of work aimed at the long tail of production inference: malformed model output, sparse attention on Qwen4, and an indexer that OOMs with room to spare in the cache. In total, 90 commits, 90 pull requests, and 4 releases landed.
Who contributed to Local LLMs on September 21, 2026?
14 developers shipped this update, including aldehir, am17an, JohannesGaessler, ggerganov, Isotr0py, jinzhen-lin, xhx1022, and yewentao256, and 6 more.
What were the notable Local LLMs updates?
common/peg : handle invalid utf-8 sequences in the AST, CUDA: enable sparse fa for qwen4, and [DeepSeek-V4.1] Bound dense prefill indexer memory.