$ cat local-llm/month/2026-09-01.log
the month in review · September 2026
vLLM 0.30.0 lands and nineteen engine-fatal flaws close
By RepoJournal · composed from the cited sources · human-reviewed weekly · methodology
The month's security arc across vLLM, SGLang, and llama.cpp ranks above every feature ship.
v0.30.0 vllm-project/vllm
The month's headline release, 762 commits from 315 contributors, lands DeepSeek-V4.1-Flash with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100, plus DeepGEMM Mega-mHC and async Engram prefetch. Upgrade for the model support, but treat the release as the fix carrier for the engine-fatal flaws vLLM disclosed the same week.
v0.29.0 vllm-project/vllm
Model Runner V2 became the default for all models here, completing a rollout that started with pooling models, with CUDA graph memory profiling added for KV cache auto-sizing. The same release added LoRA and fusion support and a max_num_queued_reqs path, since, as the PR puts it, "vLLM ships with an unbounded request queue."
Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites) vllm-project/vllm
Structured-output request errors at three sites escaped the request boundary and terminated the shared EngineCore, an engine-fatal denial of service. Any deployment accepting untrusted structured-output requests was exposed regardless of tenant isolation.
Scale-out disaggregated multimodal transport trusts caller-supplied features — shared EngineCore denial of service, encoder-cache poisoning, and transport integrity loss (5 sites) vllm-project/vllm
Five sites in the scale-out disaggregated multimodal transport trusted caller-supplied features, leaving a shared EngineCore denial of service, encoder-cache poisoning, and transport integrity loss. Distributed multimodal deployments carry the widest blast radius of the month's disclosures.
Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens in the vLLM Rust frontend metrics middleware (unauthenticated denial of service) vllm-project/vllm
Attacker-controlled HTTP method tokens fed unbounded Prometheus label cardinality in the Rust frontend metrics middleware, an unauthenticated denial of service. If your metrics scrape endpoint is reachable, cardinality is a remote kill switch until you patch.
v0.5.0 ggml-org/llama.cpp
by github-actions[bot]
The release notes say v0.5.0 "focuses on backend performance and correctness, broader model coverage, and more robust server/router operation," adding HRM-Text (DFM Mimir 1B), MiMo-V2.6 and HunyuanOCR conversion, ggml 0.25.0 improvements, and multi-address HTTP binding. It follows v0.5.19's deprecation of the --mmap, mlock, and dio flags.
v0.5.19 sgl-project/sglang
SGLang's 786 PRs from 214 contributors brought Qwen3.8 (2.4T-A95B) and a fix for the TP hangs and GLM-5.2 routing issues that surfaced mid-month. The release keeps SGLang on the same sparse-attention and disaggregation trajectory as vLLM.
[AutoRound] Support AutoRound Format Block-Wise FP8 in vLLM vllm-project/vllm
AutoRound block-wise FP8 format now loads in vLLM behind the AutoRound path, widening quantized checkpoint compatibility for anyone serving pre-quantized weights.
$ ls local-llm/month/ # the briefings behind this review
Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.
fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.
Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.