RepoJournal
Local LLMs Local LLMs
69 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-09-30
stories 298

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

Ollama gates System One on declared capabilities, llama.cpp fixes a stale-UI trap

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

Ollama moves capability decisions into the Modelfile while llama.cpp, vLLM, sglang, and MLX ship correctness and performance work underneath: a prefix-cache collision fix, a service worker that refuses to die, and FP32 Hexagon kernels.

create: support explicit model capabilities ollama/ollama

by dhiltgen

Modelfiles gain CAPABILITY declarations and create requests an additive capabilities field, preserved across GGUF and safetensors creation, inheritance, and Modelfile export. System One requests are now scheduled only when the model declares the decision capability, replacing the old Qwen architecture and renderer metadata match; GGUF-only scoring stays in place until the separate MLX runtime w...

[Core][BugFix] Tag prefix-cache extra keys by source vllm-project/vllm

by KernelClint

Prefix-cache block hashes previously mixed LoRA names, cache_salt, multimodal identifiers, and prompt-embeds digests into one untagged extra_keys tuple, so a LoRA request for adapter foo and a base-model request with cache_salt="foo" produced the same first-block hash for the same tokens. A request could then be served KV computed with another request's weights, and a caller could pick a salt l...

server : remove the built-in UI's service worker when the UI is not served (#29565) ggml-org/llama.cpp

by Emanuil Rusev

With --path or --no-ui, /sw.js returned 404, and a 404 does not unregister a service worker, so browsers kept serving the cached built-in UI. The server now returns a worker that unregisters itself, clears its caches, and reloads open tabs, while a sw.js in the --path folder still takes precedence.

[Bugfix][Responses API] Build streamed final response from streamed items vllm-project/vllm

by sfeng33

In non-harmony streaming, /v1/responses rebuilt response.completed by re-parsing the full output instead of reusing what it streamed, so the two disagreed: items got new ids, a function call's call_id changed, and clients continuing with previous_response_id referenced a call_id the stored response did not have. Streamed items could also vanish (a whitespace-only message before a tool call) and...

[Security] Bump nltk, aiohttp, pillow, and datamodel-code-generator (#59249) vllm-project/vllm

by Juan Pérez de Algaba

Upstream pins were raised for nltk, aiohttp, pillow, and datamodel-code-generator, so images and environments building vLLM from source should pick up the corrected pins rather than resolving the old ones. Those four are the security bump to take before the next build.

[Kernel][Perf] Add TP=2/4/8 per-rank shapes to the sm_120 batch-invariant matmul table vllm-project/vllm

by LioEinaudi

The sm120 batch-invariant matmul table only carried TP=1 shapes for Qwen3-1.7B/4B/8B, so VLLM_BATCH_INVARIANT=1 with tensor_parallel_size > 1 left every linear layer on RTX PRO 6000 and RTX 50 on the default 128x128x64, 8 warps, 3 stages config, leaving the roughly 2.9x batch-invariant overhead from the TP=2 rows unchanged. Per-rank TP=2/4/8 shapes are now in the table.

[Kernel] Expose FlashMLA kv_format in sgl-kernel sparse decode sgl-project/sglang

by DarkSharpness

FlashMLA's V3.2-no-RoPE and V4.1 KV caches are both 528 bytes per token at head_dim 512, so shape alone cannot separate them, and sgl-kernel always passed no format, making FlashMLA read a V4.1 cache as V3.2-no-RoPE. sparse_decode_fwd and flash_mla_with_kvcache now accept an optional kv_format.

docs: document System One API ollama/ollama

by ParthSareen

A Decision guide and System One API reference now cover choice, yes/no, and scoring questions, including local availability, request and response fields, and limits.

Quick answers

What shipped in Local LLMs on September 30, 2026?
Ollama moves capability decisions into the Modelfile while llama.cpp, vLLM, sglang, and MLX ship correctness and performance work underneath: a prefix-cache collision fix, a service worker that refuses to die, and FP32 Hexagon kernels. In total, 144 commits, 143 pull requests, and 11 releases landed.
Who contributed to Local LLMs on September 30, 2026?
20 developers shipped this update, including ParthSareen, dhiltgen, angt, aparmp-quic, Ethan Guo, Emanuil Rusev, trivikram-reddy1, and Juan Pérez de Algaba, and 12 more.
What were the notable Local LLMs updates?
create: support explicit model capabilities, [Core][BugFix] Tag prefix-cache extra keys by source, and server : remove the built-in UI's service worker when the UI is not served (#29565).