RepoJournal
Local LLMs Local LLMs
69 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-09-27
stories 136

© 2026 RepoJournal Home Showcase Explore How it works Privacy

$ the-wire · showcase

vLLM locks down multimodal kwargs, llama.cpp cleans up failed restores

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

vLLM now rejects untrusted per-request multimodal processor overrides by default, while llama.cpp closes a state-corruption path after failed sequence restores and MLX trims attention memory on NAX GPUs.

[Security] Gate per-request multimodal processor kwargs vllm-project/vllm

by jperezdealgaba

API clients could previously send per-request mm_processor_kwargs and media_io_kwargs that flowed straight into multimodal media loading and processor construction, letting untrusted overrides force extreme image or video settings and exhaust memory. Non-empty per-request kwargs are now rejected by default, with a new --trust-request-mm-kwargs flag on BaseFrontendArgs for deployments that need ...

llama : fix K/V and recurrent state cleanup after failed restores ggml-org/llama.cpp

by CHIPMUNK-T0T

A failed per-sequence state restore could delete sequence metadata while leaving K/V or recurrent-state tensor data written by that restore, and buffer-backed restores could leave deferred writes pending to be applied after cleanup. The fix clears affected tensors and discards deferred writes before the reader is destroyed, so a failed restore no longer leaves stale state behind to corrupt late...

Improve metal memory usage for SDPA D512 ml-explore/mlx

by dhiltgen

Attention at D512 on NAX GPUs previously held a larger working set than necessary; the branch cuts peak memory by 1.91 GB (7.99%) on mlx-community/gemma-4-31b-it-nvfp4 at 16K and by 4.22 GB (30.6%) on the 12B model at 65K in the PR's measurements. Larger context windows on Gemma 4-class models get headroom without changing results.

[Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights vllm-project/vllm

by Willian-Zhang

RoutedExperts.load_weights scanned the entire expert mapping for every checkpoint tensor, making matching quadratic in expert count; profiling of Qwen3.8-Flash-Next-NVFP4 on a DGX Spark attributed roughly 25 seconds of weight loading to that loop. A new _index_expert_mapping helper groups entries by the checkpoint's literal logical expert ID, so load time stops scaling with expert count squared.

cuda: add F16 input to the FWHT (#29096) ggml-org/llama.cpp

by bri-prism

The CUDA FWHT accepted F32 input only, forcing a converted copy; the source type is now a template parameter so the kernel reads F16 directly, with supports_op accepting an F16 src1 against an F32 src0 for the Hadamard hint and every other F16 src1 against a non-F16 src0 still refused. The F32 path is unchanged, and ggml_cuda_op_mul_mat_use_fwht remains the single predicate shared by supports_o...

Quick answers

What shipped in Local LLMs on September 27, 2026?
vLLM now rejects untrusted per-request multimodal processor overrides by default, while llama.cpp closes a state-corruption path after failed sequence restores and MLX trims attention memory on NAX GPUs. In total, 63 commits, 63 pull requests, and 10 releases landed.
Who contributed to Local LLMs on September 27, 2026?
13 developers shipped this update, including dhiltgen, bri-prism, CHIPMUNK-T0T, kurquhar, max-krasnyansky, Cyrus Leung, Juan Pérez de Algaba, and yewentao256, and 5 more.
What were the notable Local LLMs updates?
[Security] Gate per-request multimodal processor kwargs, llama : fix K/V and recurrent state cleanup after failed restores, and Improve metal memory usage for SDPA D512.