RepoJournal

$ cat local-llm/week/2026-09-21.log

Local LLMs

Local LLMs

the week in review · Sep 21 – Sep 27, 2026

Five engine-fatal flaws hit vLLM, llama.cpp ships v0.5.0

By RepoJournal · composed from the cited sources · human-reviewed weekly · methodology

♥

vLLM's disclosure covers request-handling paths that take down the shared EngineCore, affecting anyone serving structured output or multimodal traffic.

806 commits 799 PRs merged 67 releases 7 security advisories 7 briefings covered

all local-llm reviews →

Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites) vllm-project/vllm

Errors raised while parsing structured output cross the request boundary and terminate the shared EngineCore, so one bad request takes down every tenant on the engine. Three sites are named; if you expose guided decoding or JSON-schema output to untrusted callers, that path is now a denial-of-service vector.

Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens in the vLLM Rust frontend metrics middleware (unauthenticated denial of service) vllm-project/vllm

The Rust frontend's metrics middleware uses attacker-controlled HTTP method tokens as Prometheus labels, so unauthenticated requests inflate label cardinality without bound. Metrics backends fall over before the engine does, which makes this a cheap hit for anyone reachable from the internet.

Scale-out disaggregated multimodal transport trusts caller-supplied features — shared EngineCore denial of service, encoder-cache poisoning, and transport integrity loss (5 sites) vllm-project/vllm

The scale-out disaggregated multimodal transport trusts caller-supplied features across five sites, with the disclosure listing a shared EngineCore denial of service, encoder-cache poisoning, and transport integrity loss. This is the one to triage first if you run prefill-decode disaggregation with image or video input.

Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the num_frames ceiling does not reach vllm-project/vllm

Qwen2-VL and Qwen3-VL video samplers bound on a request-controlled max_frames value that the num_frames ceiling never reaches, so the limit you set does not actually cap the work. Video endpoints serving these models accept more frames than configured.

v0.5.0 ggml-org/llama.cpp

by github-actions[bot]

The release bundles backend performance and correctness work with HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion, multi-address HTTP binding, and image output, on top of ggml 0.25.0. It is the period's broadest single upgrade for self-hosted inference.

[Security] Reject min_tokens that exceeds the filled max_tokens default (#57731) vllm-project/vllm

by Juan Pérez de Algaba

Requests whose min_tokens exceeds the filled max_tokens default are now rejected rather than accepted into the sampling path. Worth knowing if you generate sampling parameters programmatically, since those requests will start failing instead of silently misbehaving.

common/peg : handle invalid utf-8 sequences in the AST ggml-org/llama.cpp

by aldehir

The parser now records invalid UTF-8 sequences during parse and exposes a sanitized_text() at every node in the AST, addressing models that occasionally emit bad bytes and break parsing. Tool names and arguments are deliberately excluded, on the stated grounds that they should be grammar-constrained.

server: apply structured outputs in a single pass on thinking models ollama/ollama

by jessegross

A format on a thinking model currently runs two generations: an unconstrained pass the server cancels once the parser sees content, then a re-rendered prompt under the grammar. Doing it in one pass removes the second prefill, the dropped chunk at the boundary, and the harmony prompt hack.

$ ls local-llm/week/ # the briefings behind this review

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

all local-llm reviews →