$ the-wire · showcase
MLX ports Metal gather_qmm wins to CUDA, heads-up on a qmm_sm80 corruption bug
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
MLX's CUDA MoE path got the biggest single-day speedup across the stack (Nemotron 3 Nano prompt throughput nearly doubled on an RTX 5090), while llama.cpp, vLLM, and SGLang landed targeted fixes for WebGPU aliasing, ROCm decode latency, and LoRA unload hangs.
[CUDA] Add fp_gather_qmv to optimize gather_qmm ml-explore/mlx
MLX added fp_gather_qmv on CUDA, which speeds up gather_qmm for MoE models: in mlx-lm benchmarks at p2048/g128, NVIDIA-Nemotron-3-Nano-30B-A3B went from 1671.6 to 3177.8 prompt tps on an RTX 5090 and gemma-4-26b-a4b-it from 1575.2 to 2553.0, with generation speed unchanged. The PR notes it carries a few lines from a separate change.
[CUDA] Add gather_qmm_rhs_sm80 ml-explore/mlx
gather_qmm_rhs_sm80 ports the Metal RHS optimizations to CUDA, shrinking tile size when expert groups are shorter than a 64-row tile (qwen3-coder 30B incremental +23% at p512, +6% at 8k and up). The same PR fixes a data corruption bug in qmm_sm80_mainloop, where the loop issued cp.async past the last K tile without waiting for the tail.
[ROCm][Kimi-K3][Perf] Fuse MLA decode KV-cache write and Q-prep via AITER vllm-project/vllm
On ROCm, Kimi-K3 MLA decode ran the fp8 KV-cache write and the query concat/quant as separate launches; a new AITER fused_qk_rope_concat_and_cache_mla path collapses them into one, taking a rank from 288 to 144 kernel launches and 1351.8 to 682.8 µs (1.98x) over 6 decode steps.
webgpu: fix SSM_SCAN binding aliasing ggml-org/llama.cpp
WebGPU SSM_SCAN computed its merged binding range from all four of x, dt, B, and C, but Falcon-H1 places dst between dt and x/B/C, so the merged binding overlapped dst and tripped the CI aliasing error. The fix narrows handling to the x/B, B/C, and x/B/C overlap patterns.
Fix LoRA usage-counter accounting across request lifecycle paths sgl-project/sglang
LoRARegistry counts in-flight adapter usage with acquire()/release(), and wait_for_unload() blocks until that counter is exactly zero, so a leaked acquire hangs unload forever. This fixes the accounting imbalance across request lifecycle paths.
Improve pre-M5 metal memory usage for SDPA D512 ml-explore/mlx
Pre-M5 Metal SDPA at D512, the path gemma-4 models hit, drops peak memory with no throughput trade: M4 Pro on Gemma 4 12B at 16K went from 9.736 GB to 9.032 GB while generation held at 29.410 to 29.418 tps. Elsewhere in the long tail: llama.cpp's CPU ops CSV was stale since commit 0f1e9d1 and is missing 11 supported ops, MX.split fixed its data size span on non-contiguous slices, and SGLang lan...