RepoJournal

$ cat local-llm/month/2026-09-01.log

Local LLMs

Local LLMs

the month in review · September 2026

vLLM 0.30.0 lands and nineteen engine-fatal flaws close

By RepoJournal · composed from the cited sources · human-reviewed weekly · methodology

♥

The month's security arc across vLLM, SGLang, and llama.cpp ranks above every feature ship.

3419 commits 3394 PRs merged 289 releases 21 security advisories 30 briefings covered

all local-llm reviews →

v0.30.0 vllm-project/vllm

by khluu

The month's headline release, 762 commits from 315 contributors, lands DeepSeek-V4.1-Flash with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100, plus DeepGEMM Mega-mHC and async Engram prefetch. Upgrade for the model support, but treat the release as the fix carrier for the engine-fatal flaws vLLM disclosed the same week.

v0.29.0 vllm-project/vllm

by khluu

Model Runner V2 became the default for all models here, completing a rollout that started with pooling models, with CUDA graph memory profiling added for KV cache auto-sizing. The same release added LoRA and fusion support and a max_num_queued_reqs path, since, as the PR puts it, "vLLM ships with an unbounded request queue."

Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites) vllm-project/vllm

Structured-output request errors at three sites escaped the request boundary and terminated the shared EngineCore, an engine-fatal denial of service. Any deployment accepting untrusted structured-output requests was exposed regardless of tenant isolation.

Scale-out disaggregated multimodal transport trusts caller-supplied features — shared EngineCore denial of service, encoder-cache poisoning, and transport integrity loss (5 sites) vllm-project/vllm

Five sites in the scale-out disaggregated multimodal transport trusted caller-supplied features, leaving a shared EngineCore denial of service, encoder-cache poisoning, and transport integrity loss. Distributed multimodal deployments carry the widest blast radius of the month's disclosures.

Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens in the vLLM Rust frontend metrics middleware (unauthenticated denial of service) vllm-project/vllm

Attacker-controlled HTTP method tokens fed unbounded Prometheus label cardinality in the Rust frontend metrics middleware, an unauthenticated denial of service. If your metrics scrape endpoint is reachable, cardinality is a remote kill switch until you patch.

v0.5.0 ggml-org/llama.cpp

by github-actions[bot]

The release notes say v0.5.0 "focuses on backend performance and correctness, broader model coverage, and more robust server/router operation," adding HRM-Text (DFM Mimir 1B), MiMo-V2.6 and HunyuanOCR conversion, ggml 0.25.0 improvements, and multi-address HTTP binding. It follows v0.5.19's deprecation of the --mmap, mlock, and dio flags.

v0.5.19 sgl-project/sglang

by Qiaolin-Yu

SGLang's 786 PRs from 214 contributors brought Qwen3.8 (2.4T-A95B) and a fix for the TP hangs and GLM-5.2 routing issues that surfaced mid-month. The release keeps SGLang on the same sparse-attention and disaggregation trajectory as vLLM.

[AutoRound] Support AutoRound Format Block-Wise FP8 in vLLM vllm-project/vllm

by Zhenzhong1

AutoRound block-wise FP8 format now loads in vLLM behind the AutoRound path, widening quantized checkpoint compatibility for anyone serving pre-quantized weights.

$ ls local-llm/month/ # the briefings behind this review

Tue Sep 1 vLLM bounds validation errors, AutoRound FP8 arrives Wed Sep 2 Ollama honors model generation defaults; llama.cpp and vLLM ship MoE performance fixes Thu Sep 3 Error propagation hardened in ollama's MLX bindings, gemma4 gains multimodal support Fri Sep 4 Ollama enables speculative decoding under structured output, vLLM adds LoRA and fusion Sat Sep 5 vLLM Fast Start maps weights zero-copy; GLMGA sampling capped Sun Sep 6 vLLM closes scale-out multimodal handoff hole Mon Sep 7 Vulkan adds TQ1_0, Spark2.5 lands in llama.cpp Tue Sep 8 SGLang fixes TP hangs and GLM-5.2 routing Wed Sep 9 Ollama retries compaction after context overflow, llama.cpp fuses Vulkan kernels Thu Sep 10 vLLM 0.29.0 makes Model Runner V2 the default, llama.cpp deprecates --mmap|mlock|dio Fri Sep 11 Ollama fixes unbounded KV memory growth, vLLM clarifies CVE timing Sat Sep 12 Ollama drops its CLI agent, llama.cpp cuts build times, vLLM fixes crash under DP>1 Sun Sep 13 vLLM EngineCore can be killed by crafted stop_token_ids, llama.cpp reworks JSON schema handling Mon Sep 14 SYCL memory reporting fixed, sparse-attention kernels land across vLLM and SGLang Tue Sep 15 llama.cpp 0.4.1 lands Maple 20B-A1B, Ollama drops typical_p Wed Sep 16 GPU backends converge on sparse attention as vLLM wires DeepSeek V4.1 indexer Thu Sep 17 MLX engine graduates out of x/, ROCm fixes silent attention corruption Fri Sep 18 GGUF offset corruption fixed, vLLM restores Dynamo KV metadata Sat Sep 19 Ollama exposes thinking levels, llama.cpp widens Hexagon and architecture coverage Sun Sep 20 Metal MoE/SSM fusion lands, MiMo-V2.5 fp8 sharding fixed Mon Sep 21 llama.cpp sanitizes invalid UTF-8 in the parser AST Tue Sep 22 vLLM 0.30.0 lands DeepSeek-V4.1-Flash, and a min_tokens validation hole closes Wed Sep 23 Ollama folds structured output into one pass, llama.cpp binds multiple addresses Thu Sep 24 llama.cpp v0.5.0 ships, vLLM discloses five engine-fatal request-handling flaws Fri Sep 25 Ollama eases off typical_p, llama.cpp lands Vulkan int8 tensor cores Sat Sep 26 Tiled k-quant matmul lands in llama.cpp, SYCL gets sparse FA Sun Sep 27 vLLM locks down multimodal kwargs, llama.cpp cleans up failed restores Mon Sep 28 llama.cpp unblocks chunked reranking, vLLM patches three engine-core holes Tue Sep 29 vLLM's rejection path desyncs multimodal caches Wed Sep 30 Ollama gates System One on declared capabilities, llama.cpp fixes a stale-UI trap

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

all local-llm reviews →