RepoJournal
Local LLMs Local LLMs
76 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-10-06
stories 292

© 2026 RepoJournal Home Showcase How it works Privacy

$ the-wire · showcase

vLLM's request-controlled caches, and Vulkan sparse attention for quantized K/V

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

Two remote-triggerable resource bugs land in vLLM's queue while llama.cpp turns on sparse flash attention for quantized KV caches, the fix that matters most to anyone running long-context inference on non-NVIDIA GPUs.

MOSS-Audio request-controlled processor cache permits remote memory exhaustion vllm-project/vllm

A request-controlled processor cache in MOSS-Audio lets a remote caller exhaust memory, so a single crafted request is enough to take a serving instance down. Treat this as the day's most urgent item if you expose MOSS-Audio endpoints to untrusted traffic.

An oversized min_tokens in one request hangs the vLLM engine while /health stays green vllm-project/vllm

An oversized min_tokens in one request hangs the engine while /health keeps reporting green, which means your orchestrator will not restart the pod for you. Until a fix lands, cap min_tokens at the gateway rather than trusting readiness probes.

vulkan: sparse flash attention for quantized K/V ggml-org/llama.cpp

by fxgsell

Vulkan's sparse flash attention previously activated only when K and V were f16, so quantized-cache models like Qwen3-Flash-Next QSA and DeepSeek-style sparse attention ran dense attention over the entire context despite keeping only about 2k cells. The scalar and coopmat1 paths already dequantize each gathered row through the sparse index, so the gate just needed to admit quantized K/V; coopma...

mlx: mitigate high latency after GPU idle ollama/ollama

by dhiltgen

ollama now carries the MLX residency-refresh patch and enables it every second in the runner, preserving explicit environment overrides, which targets the high-latency stall after the GPU has sat idle. The same author also landed model-lookup and MLX decision-request overhead cuts, avoiding decodes of unrelated manifests and reusing small Metal scratch buffers between decision requests.

[HiCache] Route --file-storage-path to the file storage backend sgl-project/sglang

by reger-men

--file-storage-path is parsed into server_args.file_storage_path but nothing reads it: the file HiCache backend honors only SGLANG_HICACHE_FILE_BACKEND_STORAGE_DIR and otherwise writes to /tmp/hicache, so L3 silently lands on the wrong disk or on a tmpfs that is far smaller than intended.

Fix metal validation error when passing empty vector ml-explore/mlx

by aleroot

Elsewhere: MLX replaced Black with ruff, valid scalar and zero-index GPU ops got typed placeholder bindings to stop Metal validation aborting Gather::eval_gpu, and default.metallib lookup now resolves under swift test with the Swift Build system, a case where MLX previously killed the entire test process on the first GPU call.

Quick answers

What shipped in Local LLMs on October 6, 2026?
Two remote-triggerable resource bugs land in vLLM's queue while llama.cpp turns on sparse flash attention for quantized KV caches, the fix that matters most to anyone running long-context inference on non-NVIDIA GPUs. In total, 141 commits, 139 pull requests, 10 releases, and 2 security advisories landed.
Who contributed to Local LLMs on October 6, 2026?
18 developers shipped this update, including dhiltgen, dongluochen, fxgsell, ggerganov, eapache, alanhc, ngxson, and David Cheung, and 10 more.
What were the notable Local LLMs updates?
MOSS-Audio request-controlled processor cache permits remote memory exhaustion, An oversized min_tokens in one request hangs the vLLM engine while /health stays green, and vulkan: sparse flash attention for quantized K/V.