RepoJournal
Local LLMs Local LLMs
69 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-09-23
stories 288

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

Ollama folds structured output into one pass, llama.cpp binds multiple addresses

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

Ollama cut the double-prefill cost of JSON-constrained thinking models and sped up Qwen prompt processing on MLX, while llama.cpp made llama-server bindable to several addresses at once and shaved the iq1_m quantizer from O(B³) to O(B²).

server: apply structured outputs in a single pass on thinking models ollama/ollama

by jessegross

A format on a thinking model used to run two generations: an unconstrained one the server cancelled once the parser reported content, then a re-rendered prompt under the grammar. Parsers now report the strings that end a response, so the same request completes in a single pass, which removes the restart's second prefill, the chunk dropped at the boundary, and a stray first token that could leak...

mlx: speed up Qwen 3.8 prompt processing ollama/ollama

by dhiltgen

dhiltgen's MLX path uses the gated-delta kernel for long scans and folds dense MLP global scales into SwiGLU, lifting prompt processing on an M5 Max from 715.2 to 848.5 tokens/s at 2k context (+18.7%), 695.6 to 828.1 at 8k, and 704.6 to 802.6 at 16k.

server: Add support for binding to multiple addresses ggml-org/llama.cpp

by erusev

Binding llama-server to a VPN address previously made it unreachable through localhost, breaking clients that default there. --host now takes a comma-separated list of IP addresses and Unix socket paths, one listener per address sharing the same HTTP worker pool, so a single server answers on both; with --port 0 the first TCP listener selects the port.

ggml-quant: Quant creation speed, IQ1_M build prefix sums once per block ggml-org/llama.cpp

by bartowski1182

iq1_m recomputed weighted sums across all 16 elements for every candidate; the quantizer now builds prefix sums once per block and scores in constant time, dropping runtime from O(B³) to O(B²). The iq1_s quantizer already did this, and float sum ordering means the output is not byte-identical.

mlx: select Gemma 4 image resolution dynamically ollama/ollama

by dhiltgen

The MLX Gemma 4 path replaces the fixed checkpoint image budget with per-image selection across the 70, 140, 280, 560, and 1120 budgets, picking the publisher resize grid closest to the input resolution with aspect ratio accounted for. High-resolution documents keep more detail and smaller images spend fewer tokens, with no new API parameter.

[dsv4.1]Optimize FP4 indexer by skipping invisible tiles sgl-project/sglang

by shiyu7

sglang's legacy FP4 indexer replayed a capacity-sized CUDA Graph grid even on short contexts: in a measured shape with 1,048,896 output columns and 8,704 visible positions, over 99% of columns were invisible yet still ran dequantization, dot products, weighting, and reduction. The indexer now skips arithmetic for fully invisible tiles and writes -inf directly to valid output addresses.

[HiCache] Remove the unused HiRadixCache sgl-project/sglang

by hnyls2002

Elsewhere on the long tail, sglang deleted the unused HiRadixCache along with its helpers and tests, since HiCache has gone through UnifiedRadixCache since v0.5.19, and dropped the deprecated SGLANG_ENABLE_UNIFIED_RADIX_TREE env now that the unified radix tree is the default; vllm removed its AllSpark INT8 W8A16 GEMM backend and dropped the model_runner_v2 test from Intel GPU CI, and llama.cpp ...

Quick answers

What shipped in Local LLMs on September 23, 2026?
Ollama cut the double-prefill cost of JSON-constrained thinking models and sped up Qwen prompt processing on MLX, while llama.cpp made llama-server bindable to several addresses at once and shaved the iq1_m quantizer from O(B³) to O(B²). In total, 140 commits, 137 pull requests, and 11 releases landed.
Who contributed to Local LLMs on September 23, 2026?
17 developers shipped this update, including ParthSareen, dhiltgen, jessegross, bartowski1182, ngxson, erusev, fish-jiang, and AesSedai, and 9 more.
What were the notable Local LLMs updates?
server: apply structured outputs in a single pass on thinking models, mlx: speed up Qwen 3.8 prompt processing, and server: Add support for binding to multiple addresses.