RepoJournal
Local LLMs Local LLMs
69 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-09-28
stories 221

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

llama.cpp unblocks chunked reranking, vLLM patches three engine-core holes

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

A llama.cpp server fix lets causal rerankers finally use chunked prefill, while vLLM closed three engine-core bugs, one a cross-tenant cache oracle and one a native heap overflow in the transcription path.

server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876) ggml-org/llama.cpp

by Tim Wang

Rerank models split into two families: bidirectional cross-encoders (BERT and friends) that need every token in one physical batch, and causal LLMs repurposed as rerankers such as Qwen3 and Qwen3-VL, which can use chunked prefill like any decoder. The server used to reject every RANK-pooling input larger than n_ubatch; it now splits them for the causal case, so large rerank inputs no longer nee...

Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle vllm-project/vllm

Harmony tool continuations stopped carrying cache_salt, which restored a cross-tenant prefix-cache membership oracle. If you serve multiple tenants behind one engine with Harmony-style tool calls, treat this as a breaking fix and upgrade before you keep sharing a cache across trust boundaries.

Crafted IAMF audio upload reaches a PyAV/FFmpeg native heap overflow through the speech transcription path — denial of service vllm-project/vllm

A crafted IAMF audio upload reaches a native heap overflow in PyAV/FFmpeg through the speech transcription path, giving a denial of service. Anything accepting untrusted audio for transcription is exposed until you take the fixed build.

[Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684) vllm-project/vllm

by Nicolò Lucchesi

FlashInfer's SM90 sparse MLA plan no longer forces a device-to-host sync when async scheduling is on, removing a stall that sat on the critical path of every plan. This is a perf change, not a behavior one: same outputs, less host round-tripping.

Add Metal SDPA support for D96/V64 ml-explore/mlx

by wyanzhao

MLX added Metal scaled dot-product attention for Dqk=96, Dv=64, the shape MiniCPM3 uses; the one-pass and two-pass vector kernels now cover it, with separate query/key and value tile widths in the NAX and ordinary full-attention kernels, and the ordinary kernel handling D96/V64 float32 with TF32 disabled. JIT and no-JIT non-NAX builds each pass 28 SDPA tests with 3 platform skips.

Add fused Metal kernels for fast.cross_entropy ml-explore/mlx

by zsun6

mx.fast.cross_entropy gained fused Metal kernels modeled on the existing CUDA ones; the old path fell back to logsumexp - take_along_axis, did the lse - x_t subtraction in the logits dtype and only then cast to float32, losing bits with bf16 logits. Against float64 numpy on random (4, 7, 8192) logits, max abs error drops from 1.0e-01 to 9.3e-07 in bfloat16.

Quick answers

What shipped in Local LLMs on September 28, 2026?
A llama.cpp server fix lets causal rerankers finally use chunked prefill, while vLLM closed three engine-core bugs, one a cross-tenant cache oracle and one a native heap overflow in the transcription path. In total, 104 commits, 104 pull requests, 10 releases, and 3 security advisories landed.
Who contributed to Local LLMs on September 28, 2026?
12 developers shipped this update, including Tim Wang, bri-prism, Adrien Gallouët, Aman Gupta, Ruben Ortlam, Nicolò Lucchesi, Turner Jabbour, and Cheng Wan, and 4 more.
What were the notable Local LLMs updates?
server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876), Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle, and Crafted IAMF audio upload reaches a PyAV/FFmpeg native heap overflow through the speech transcription path — denial of service.