$ the-wire · showcase
Tiled k-quant matmul lands in llama.cpp, SYCL gets sparse FA
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
llama.cpp's CPU matmul rewrite and a new SYCL sparse-attention path headline a day of targeted kernel work across llama.cpp, vLLM, SGLang, and MLX.
ggml-cpu: tiled mul_mat for k-quants (#27851) ggml-org/llama.cpp
ggml-cpu now unpacks k-quants into 256x256 int8 tiles and runs a 16x16 microkernel before writing float results back to main memory, with tests in tests/test-tiled-mulmat.cpp. jbooth reports 3-6x faster large matmul, break-even at 4096x64 * 64x4096, an 80 percent net loss for GEMV, and error rates on the order of 1e-04 max and 1e-05 RMSE: a workload-shaped tradeoff, not a free win.
[SYCL] support sparse FA ggml-org/llama.cpp
Intel Arc users get an opt-in sparse FlashAttention path for long-context decode, aimed at qwen4exp on the B70 and contributed by logari81. It is disabled by default and gated on GGML_SYCL_SPARSE_FA, GGML_SYCL_SPARSE_FA_DEBUG, and GGML_SYCL_SPARSE_FA_MARGIN=256, so you must export GGML_SYCL_SPARSE_FA=1 to enable it.
Improve metal memory usage for SDPA D256 ml-explore/mlx
dhiltgen's change cuts Metal memory for D256 attention on M4 and older GPUs, which is where quantized Qwen checkpoints were running out of headroom, and adds a small prompt speedup. Testing with mlx-community/Qwen3.8-27B-nvfp4 and mlx-community/Qwen3.5-2B-nvfp4 on an M4 Pro with 64 GB shows the effect at 16K prompt length.
[Core][Logging] Fix JSON logging process decoration vllm-project/vllm
decorate_logs() was wrapping stdout and stderr to prefix output with the logical process name and PID, which corrupted JSON from custom Python logging configs, including the documented python-json-logger example: every object gained non-JSON leading text. The name now lives in a LogRecord attribute, and the default text formatter renders it alongside the standard PID, so JSON consumers get clea...
[LMCache] Support lmcache unified radix cache sgl-project/sglang
LMCache can now serve as an external KV-cache backend on the UnifiedRadixCache path, persisting and restoring reusable KV blocks through it while keeping the existing radix-cache lookup, store, reuse, and flush behavior intact. Also worth noting today: a llama.cpp RPC fix that adds nb to the get_alloc_size cache key and floors the result at ggml_nbytes, vec and register-resident FP8 quantizatio...