Local LLMs
Vulkan adds TQ1_0, Spark2.5 lands in llama.cpp
llama.cpp and its ecosystem shipped new hardware support and a model architecture, while vLLM and SGLang fixed bugs in prefill and warmup paths.
read --wire →
$ tail -f topics/llms.log
Daily updates from the local-LLM inference stack - Ollama, llama.cpp, vLLM, and SGLang. What shipped this week in running and serving open models on your own hardware.
11 updates across 2 projects this week.
One calm review of what shipped across Local LLMs - the commits, releases, and security advisories that matter. Every Monday, with security advisories same-day. Free, unsubscribe in one click.
We'll start you on the top Local LLMs projects - refine anytime. · Read a sample issue →
Local LLMs
llama.cpp and its ecosystem shipped new hardware support and a model architecture, while vLLM and SGLang fixed bugs in prefill and warmup paths.
read --wire →
Open WebUI
A fix removes asyncio.wait_for/shield from MCP cleanup in the chat handler, preventing a BaseException crash on anyio cancel-scope violations.
read --wire →
Open WebUI
Open WebUI Desktop 0.0.19 ships with a critical Linux fix that restores blank webviews by replacing --in-process-gpu with SwiftShader.
read --wire →
Open WebUI
Open WebUI Desktop shipped three fixes that repair broken terminal startup, blank webview content on Linux, and hidden Hugging Face models.
read --wire →
Open WebUI
Open WebUI desktop release 0.0.17 fixes external links so they open in the default browser instead of inside the app.
read --wire →
Local LLMs
vLLM now validates scale-out multimodal data before engine handoff, closing a trust gap in the generate route.
read --wire →
Local LLMs
Two security patches and a zero-copy weight cache land across vLLM and llama.cpp, while sglang ships v0.5.19 with 786 PRs.
read --wire →
Local LLMs
Ollama's MLX runner now runs speculative decoding under structured output constraints, promising near-full draft throughput on MLX models, while vLLM lands static LoRA loading in its Rust frontend and begins manual activation-quant fusion.
read --wire →
Local LLMs
ollama's MLX runner now treats every fallible MLX call as an error it can surface, ending silent failures that corrupted outputs.
read --wire →
Local LLMs
Ollama now respects model-authored sampler defaults from GGUF metadata, and both llama.cpp and vLLM land performance and correctness fixes for MoE kernels.
read --wire →
Local LLMs
vLLM shipped a security fix bounding validation-error response bodies and added AutoRound block-wise FP8 support, while llama.cpp and sglang pushed kernel fusions for Blackwell and ROCm.
read --wire →
Local LLMs
Across three repos, the most consequential changes tune performance on AMD and Vulkan, fix a vLLM memory regression, and extend speculative decoding support.
read --wire →
Local LLMs
llama.cpp fixes a Vulkan memory bloat bug that caused extreme VRAM usage, while vLLM begins deprecating its PyAV video decoder backend over performance concerns.
read --wire →
Local LLMs
llama.cpp's OpenVINO backend now runs Qwen3.5 on Intel NPUs and supports Whisper.cpp, while sglang fixes a memory threshold that was causing Cosmos3 Nano to offload on 96 GB GPUs.
read --wire →
Local LLMs
The biggest local-model news in weeks: llama.cpp just landed draft support for Qwen3.8-Flash-Next, and it ships with three quantizer fixes the model needs.
read --wire →
Local LLMs
vLLM shipped its biggest release in months, and Ollama quietly fixed a macOS bug that could leave stale app processes running.
read --wire →
Local LLMs
Ollama closed a silent correctness gap on Apple Silicon, and the LLM inference stack shipped a wave of fixes that touch everything from tool calls to native ARM support.
read --wire →
Local LLMs
Ollama's desktop app just made Claude Desktop integration faster and safer to manage, while vLLM closed a security gap on audio file size limits.
read --wire →
Local LLMs
Forget all-gathering experts during RL rollout: vLLM's new sharded_rdt backend lets each worker pull only the slice it needs, turning a total_bytes transfer into a total_bytes divided by num_workers one.
read --wire →
Local LLMs
The biggest news overnight: SGLang is patching a production-breaking Pixtral bug, and vLLM engineers found a way to shave up to 25% off time-to-first-token on Mamba models.
read --wire →