RepoJournal
Local LLMs Local LLMs
70 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-10-01
stories 330

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

MLX ports Metal gather_qmm wins to CUDA, heads-up on a qmm_sm80 corruption bug

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

MLX's CUDA MoE path got the biggest single-day speedup across the stack (Nemotron 3 Nano prompt throughput nearly doubled on an RTX 5090), while llama.cpp, vLLM, and SGLang landed targeted fixes for WebGPU aliasing, ROCm decode latency, and LoRA unload hangs.

[CUDA] Add fp_gather_qmv to optimize gather_qmm ml-explore/mlx

by dhiltgen

MLX added fp_gather_qmv on CUDA, which speeds up gather_qmm for MoE models: in mlx-lm benchmarks at p2048/g128, NVIDIA-Nemotron-3-Nano-30B-A3B went from 1671.6 to 3177.8 prompt tps on an RTX 5090 and gemma-4-26b-a4b-it from 1575.2 to 2553.0, with generation speed unchanged. The PR notes it carries a few lines from a separate change.

[CUDA] Add gather_qmm_rhs_sm80 ml-explore/mlx

by dhiltgen

gather_qmm_rhs_sm80 ports the Metal RHS optimizations to CUDA, shrinking tile size when expert groups are shorter than a 64-row tile (qwen3-coder 30B incremental +23% at p512, +6% at 8k and up). The same PR fixes a data corruption bug in qmm_sm80_mainloop, where the loop issued cp.async past the last K tile without waiting for the tail.

[ROCm][Kimi-K3][Perf] Fuse MLA decode KV-cache write and Q-prep via AITER vllm-project/vllm

by rbrugaro-amd

On ROCm, Kimi-K3 MLA decode ran the fp8 KV-cache write and the query concat/quant as separate launches; a new AITER fused_qk_rope_concat_and_cache_mla path collapses them into one, taking a rank from 288 to 144 kernel launches and 1351.8 to 682.8 µs (1.98x) over 6 decode steps.

webgpu: fix SSM_SCAN binding aliasing ggml-org/llama.cpp

by yomaytk

WebGPU SSM_SCAN computed its merged binding range from all four of x, dt, B, and C, but Falcon-H1 places dst between dt and x/B/C, so the merged binding overlapped dst and tripped the CI aliasing error. The fix narrows handling to the x/B, B/C, and x/B/C overlap patterns.

Fix LoRA usage-counter accounting across request lifecycle paths sgl-project/sglang

by shafeeqibraheem

LoRARegistry counts in-flight adapter usage with acquire()/release(), and wait_for_unload() blocks until that counter is exactly zero, so a leaked acquire hangs unload forever. This fixes the accounting imbalance across request lifecycle paths.

Improve pre-M5 metal memory usage for SDPA D512 ml-explore/mlx

by dhiltgen

Pre-M5 Metal SDPA at D512, the path gemma-4 models hit, drops peak memory with no throughput trade: M4 Pro on Gemma 4 12B at 16K went from 9.736 GB to 9.032 GB while generation held at 29.410 to 29.418 tps. Elsewhere in the long tail: llama.cpp's CPU ops CSV was stale since commit 0f1e9d1 and is missing 11 supported ops, MX.split fixed its data size span on non-contiguous slices, and SGLang lan...

Quick answers

What shipped in Local LLMs on October 1, 2026?
MLX's CUDA MoE path got the biggest single-day speedup across the stack (Nemotron 3 Nano prompt throughput nearly doubled on an RTX 5090), while llama.cpp, vLLM, and SGLang landed targeted fixes for WebGPU aliasing, ROCm decode latency, and LoRA unload hangs. In total, 160 commits, 160 pull requests, and 10 releases landed.
Who contributed to Local LLMs on October 1, 2026?
18 developers shipped this update, including CaramelizedCUDA, Vishal Singh, yomaytk, ebateni, Hrishith Thadicherla, rbrugaro-amd, benchislett, and TQCB, and 10 more.
What were the notable Local LLMs updates?
[CUDA] Add fp_gather_qmv to optimize gather_qmm, [CUDA] Add gather_qmm_rhs_sm80, and [ROCm][Kimi-K3][Perf] Fuse MLA decode KV-cache write and Q-prep via AITER.