RepoJournal
Local LLMs Local LLMs
69 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-09-26
stories 252

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

Tiled k-quant matmul lands in llama.cpp, SYCL gets sparse FA

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

llama.cpp's CPU matmul rewrite and a new SYCL sparse-attention path headline a day of targeted kernel work across llama.cpp, vLLM, SGLang, and MLX.

ggml-cpu: tiled mul_mat for k-quants (#27851) ggml-org/llama.cpp

by jbooth

ggml-cpu now unpacks k-quants into 256x256 int8 tiles and runs a 16x16 microkernel before writing float results back to main memory, with tests in tests/test-tiled-mulmat.cpp. jbooth reports 3-6x faster large matmul, break-even at 4096x64 * 64x4096, an 80 percent net loss for GEMV, and error rates on the order of 1e-04 max and 1e-05 RMSE: a workload-shaped tradeoff, not a free win.

[SYCL] support sparse FA ggml-org/llama.cpp

by arthw

Intel Arc users get an opt-in sparse FlashAttention path for long-context decode, aimed at qwen4exp on the B70 and contributed by logari81. It is disabled by default and gated on GGML_SYCL_SPARSE_FA, GGML_SYCL_SPARSE_FA_DEBUG, and GGML_SYCL_SPARSE_FA_MARGIN=256, so you must export GGML_SYCL_SPARSE_FA=1 to enable it.

Improve metal memory usage for SDPA D256 ml-explore/mlx

by dhiltgen

dhiltgen's change cuts Metal memory for D256 attention on M4 and older GPUs, which is where quantized Qwen checkpoints were running out of headroom, and adds a small prompt speedup. Testing with mlx-community/Qwen3.8-27B-nvfp4 and mlx-community/Qwen3.5-2B-nvfp4 on an M4 Pro with 64 GB shows the effect at 16K prompt length.

[Core][Logging] Fix JSON logging process decoration vllm-project/vllm

by markmc

decorate_logs() was wrapping stdout and stderr to prefix output with the logical process name and PID, which corrupted JSON from custom Python logging configs, including the documented python-json-logger example: every object gained non-JSON leading text. The name now lives in a LogRecord attribute, and the default text formatter renders it alongside the standard PID, so JSON consumers get clea...

[LMCache] Support lmcache unified radix cache sgl-project/sglang

by chunxiaozheng

LMCache can now serve as an external KV-cache backend on the UnifiedRadixCache path, persisting and restoring reusable KV blocks through it while keeping the existing radix-cache lookup, store, reuse, and flush behavior intact. Also worth noting today: a llama.cpp RPC fix that adds nb to the get_alloc_size cache key and floors the result at ggml_nbytes, vec and register-resident FP8 quantizatio...

Quick answers

What shipped in Local LLMs on September 26, 2026?
llama.cpp's CPU matmul rewrite and a new SYCL sparse-attention path headline a day of targeted kernel work across llama.cpp, vLLM, SGLang, and MLX. In total, 121 commits, 121 pull requests, and 10 releases landed.
Who contributed to Local LLMs on September 26, 2026?
16 developers shipped this update, including R0CKSTAR, Jess Sullivan, Sigbjørn Skjæret, jbooth, arthw, mmastrac, sheralskumar, and markmc, and 8 more.
What were the notable Local LLMs updates?
ggml-cpu: tiled mul_mat for k-quants (#27851), [SYCL] support sparse FA, and Improve metal memory usage for SDPA D256.