$ cat local-llm/week/2026-09-28.log
the week in review · Sep 28 – Oct 4, 2026
vLLM fixes cache_salt oracle and multimodal IPC desync
By RepoJournal · composed from the cited sources · human-reviewed weekly · methodology
Three breaking vLLM items outrank the week's features: a cross-tenant cache oracle, an IAMF heap overflow, and a desyncing multimodal cache. SGLang ships v0.5.21 with 779 PRs.
Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle vllm-project/vllm
Tool continuations under Harmony no longer carry the salt, which the story says restores a cross-tenant prefix-cache membership oracle: an untagged cache key lets one tenant probe whether another's prefix is resident. If you serve multiple tenants from one prefix cache, treat this as a security fix, not a refactor.
Crafted IAMF audio upload reaches a PyAV/FFmpeg native heap overflow through the speech transcription path — denial of service vllm-project/vllm
A crafted IAMF audio upload passes through the speech transcription path into a native heap overflow in PyAV/FFmpeg, giving a denial of service. Untrusted audio against your transcription endpoint is the trigger path, so patch before you expose it.
Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core vllm-project/vllm
After a request is rejected, the mirrored multimodal IPC caches drift out of sync; the next request that reuses the same media hash trips a receiver assertion in the engine core. The failure needs a rejected request first, which makes it easy to miss until it takes an engine down in production.
v0.5.21 sgl-project/sglang
The release covers 779 PRs from 227 contributors and adds models including DeepSeek-V4.1 Flash and GigaChat 3.5 with cookbook entries. The tag is flagged breaking, so read the notes before you move.
create: support explicit model capabilities ollama/ollama
Modelfiles gain CAPABILITY declarations and create requests get an additive capabilities field, preserved across GGUF and safetensors creation, inheritance, and Modelfile export. System One requests now require the decision capability rather than matching on Qwen architecture, so scheduled System One calls need the declaration or they will not run.
server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876) ggml-org/llama.cpp
The server can now split RANK pooling batches for causal LLM rerankers such as Qwen3 and Qwen3-VL, which the pull request notes previously had to fit all tokens in one physical batch. Rerank throughput on those models should scale with batch splitting instead of one oversized batch.
[Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684) vllm-project/vllm
The change removes a device-to-host sync from the FlashInfer SM90 sparse MLA plan under async scheduling, so the plan no longer stalls the scheduler. It targets SM90 sparse attention specifically; other paths are untouched.
[Core][BugFix] Tag prefix-cache extra keys by source vllm-project/vllm
Prefix-cache block hashes previously mixed LoRA names, cache_salt, multimodal identifiers, and prompt-embeds digests into one untagged tuple, so a LoRA request for adapter foo and a base-model request with cache_salt="foo" could hash identically. Tagging extra keys by source keeps those lookups distinct.
$ ls local-llm/week/ # the briefings behind this review
Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.
fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.
Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.