$ the-wire · showcase
vLLM's rejection path desyncs multimodal caches
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
The day's three vLLM security reports land on the request-rejection and tool-continuation paths, while llama.cpp makes multimodal embeddings a first-class endpoint and Ollama opens a scoring API.
Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core vllm-project/vllm
When a multimodal request is rejected, the mirrored IPC caches on either side of the engine core fall out of sync; the next request that reuses the same media hash reaches a receiver assertion in the engine core. If you serve audio or images with prefix or media caching enabled, a single rejected request can plant a failure that surfaces later under a different caller's traffic.
Crafted IAMF audio upload reaches a PyAV/FFmpeg native heap overflow through the speech transcription path — denial of service vllm-project/vllm
A crafted IAMF audio upload traversed the speech transcription path into a native heap overflow in PyAV/FFmpeg, giving denial of service. Treat the transcription endpoint as untrusted input until the fix lands.
Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle vllm-project/vllm
Harmony tool continuations drop `cache_salt`, which reopens a cross-tenant prefix-cache membership oracle: a caller can probe whether another tenant's prefix is resident. Multi-tenant deployments relying on `cache_salt` for isolation should not assume it survives a tool call.
server : support typed content (vision/audio/video) input for /v1/embeddings endpoint ggml-org/llama.cpp
llama.cpp's `/v1/embeddings` endpoint now accepts typed vision, audio, and video content natively, matching the OpenAI-style API that Openrouter and other providers use for multimodal embeddings; the previous legacy API is still supported. Multimodal embedding work that relied on workarounds can move to the standard request shape.
feat: add System One scoring API ollama/ollama
Ollama added `POST /v1/systemone` for structured decisions over local Nimble and Tev models, returning a `choice` with per-option probabilities, a `noul` probability, and a `score` across an ordered rubric. Answers come from next-token scores normalized over the allowed choices, and the release notes warn that `confidence` "measures how concentrated that distribution is" rather than calibrated ...
[SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts sgl-project/sglang
Quiet elsewhere: SGLang added optional FlashInfer PCIe-IPC all-reduce for switch-free SM120 hosts, where `CustomAllreduce`, `QuickAllReduce`, and `pymscclpp` all assume NVLink, multicast, or their own fabric and every per-layer reduction fell back to NCCL; MLX sped up Metal `gather_qmm` by up to 1.32x at the sizes benchmarked on qwen3.5-397b-a17b, and llama.cpp's oneAPI CI moved to toolkit 2026.1.
Action items
- → Upgrade vLLM past the IAMF/PyAV heap overflow fix before accepting untrusted audio on speech transcription endpoints vllm-project/vllm [immediate]
- → Apply the vLLM fix that restores cache_salt on Harmony tool continuations if you run multi-tenant prefix caching vllm-project/vllm [plan]