RepoJournal
Local LLMs Local LLMs
76 wires and counting

$ follow Local LLMs

Keep up with Local LLMs in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-10-07
stories 290

© 2026 RepoJournal Home Showcase How it works Privacy

$ the-wire · showcase

vLLM's multimodal input path is remote code execution, and MLX's sort is silently wrong

By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology

Five unpatched vLLM vulnerabilities let unauthenticated callers crash or execute code through the multimodal processor path, while MLX fixes two silent-wrong-answer bugs and Ollama ships multimodal embeddings.

Request-controlled mm_processor_kwargs.code_revision allows remote code execution vllm-project/vllm

A request-controlled mm_processor_kwargs.code_revision value reaches model-loading code, so any caller who can hit the OpenAI-compatible endpoint can execute arbitrary code on the server. The same input surface also produces remote crashes and cache aliasing (a forged multimodal UUID plus processor-cache eviction drift takes down the whole engine; the multimodal EXIF hash and prefix-cache extra...

Collapse the non sorted axes to pick the contiguous sort kernel ml-explore/mlx

by kapellirohith

mx.sort and mx.argsort returned silently wrong results on transposed views on GPU, with no error and no crash, because non-sorted axes were not collapsed onto the contiguous kernel. The reproducer (a 3x4x8 array through mx.swapaxes then sort along axis=-1) compared False before and True after. If you sort or argsort non-contiguous views, the previous results you based anything on are not trustw...

llama: fix clef head reads past 2GiB on windows (#18777) ollama/ollama

by Gigrise

The Clef head read its weights with std::ifstream::seekg, and on the Windows llama-server build (libc++ 18.1.8) that seek truncates offsets to 32 bits, so clef-flash's clef.* tensors at about 9.5GB loaded backbone bytes as head weights and every /v1/systemone request failed with "Clef: non-finite logit". clef:27b was unaffected because its head and output.weight stay under 2GB; the read now goe...

model: add multimodal embeddings ollama/ollama

by pdevine

Ollama implemented the EmbeddingGemma2Model architecture on the MLX runner: a 24-layer bidirectional text encoder with PLE, shared gemma4 vision and audio towers, and mean-pool plus L2 output. /api/embed now accepts per-item media through input dicts, so one embed call can mix text and image or audio items rather than requiring separate requests.

Fix float16 overflow in clip_grad_norm ml-explore/mlx

by ishtihoss

clip_grad_norm squared and summed gradients in their own dtype, so any fp16 gradient above roughly 256 overflowed to inf, making the total norm inf, the normalizer zero, and every gradient silently zero; the reporter confirmed an fp16 value of 400 returning inf on the released 0.32.3 core. The fix sums fp16 squares in fp32 and scales in fp32 before casting back, keeping clipped gradients in the...

RPC: add `-sm tensor` ggml-org/llama.cpp

by am17an

RPC gains -sm tensor, tested across 2x Sparks connected over RDMA, which lets tensor parallelism span machines instead of confining each GPU to its own layer slice. It needs async graph_compute, a custom all_reduce, a graph uid cache like CUDA's, and set_tensor_2d/get_tensor_2d in the RPC layer, so it is a reviewable proposal more than a drop-in switch.

[DeepSeek-V4.1] Unify mHC into one state machine and drop medium-batch fusion variants sgl-project/sglang

by DarkSharpness

DeepSeek-V4.1's manifold-constrained hyper-connections now live in one module, deepseek_v4_mhc.py, with a single state machine for cross-layer state and one dispatch point per mHC step, replacing contextvar hooks threaded through unrelated code and dropping the medium-batch fusion variants. The PR is marked as generated by Claude, which is worth knowing before you read the design notes.

Quick answers

What shipped in Local LLMs on October 7, 2026?
Five unpatched vLLM vulnerabilities let unauthenticated callers crash or execute code through the multimodal processor path, while MLX fixes two silent-wrong-answer bugs and Ollama ships multimodal embeddings. In total, 137 commits, 137 pull requests, 10 releases, and 6 security advisories landed.
Who contributed to Local LLMs on October 7, 2026?
17 developers shipped this update, including drifkin, pdevine, Gigrise, ngxson, ggerganov, am17an, Kartik Gulia, and weimin023, and 9 more.
What were the notable Local LLMs updates?
Request-controlled mm_processor_kwargs.code_revision allows remote code execution, Collapse the non sorted axes to pick the contiguous sort kernel, and llama: fix clef head reads past 2GiB on windows (#18777).