$ ls local-llm/reviews/ # the step-back reads
Local LLMs in review
One review per week (and month): what shipped, what broke, what to act on.
Sep 14–20, 2026 weekly [act] 8 stories
Sparse attention lands across vLLM, SGLang, llama.cpp
Vulkan and SYCL backends got sparse Flash Attention and memory fixes, while SGLang gated Responses persistence off by default.
Sep 7–13, 2026 weekly [act] 7 stories
vLLM EngineCore can be killed by crafted stop_token_ids
Two vLLM reports this week show out-of-range stop_token_ids with min_tokens terminating EngineCore, including via the Rust HTTP/gRPC path.
Aug 31–Sep 6, 2026 weekly [act]
Vulkan top_k radix select arrives for long rows
llama.cpp's Vulkan backend adds top_k radix select; vLLM security patch bounds validation-error responses, and AutoRound FP8 support ships.
Aug 24–30, 2026 weekly
vLLM 0.28.0 lands with P2P weight sync, llama.cpp and partners broaden model reach
vLLM 0.28.0 shipped with 584 commits, headlined by a peer-to-peer weight synchronization for RL workloads, while llama.cpp, Ollama, and S...
Aug 1–31, 2026 monthly
Speculative decoding consolidates across the stack as DeepSeek V4 lands
August 2026 saw the local LLM stack converge on speculative decoding and DeepSeek V4, with every major engine shipping optimized paths.
Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.
fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.
Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.