$ the-wire · showcase
vLLM's multimodal input path is remote code execution, and MLX's sort is silently wrong
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
Five unpatched vLLM vulnerabilities let unauthenticated callers crash or execute code through the multimodal processor path, while MLX fixes two silent-wrong-answer bugs and Ollama ships multimodal embeddings.
Request-controlled mm_processor_kwargs.code_revision allows remote code execution vllm-project/vllm
A request-controlled mm_processor_kwargs.code_revision value reaches model-loading code, so any caller who can hit the OpenAI-compatible endpoint can execute arbitrary code on the server. The same input surface also produces remote crashes and cache aliasing (a forged multimodal UUID plus processor-cache eviction drift takes down the whole engine; the multimodal EXIF hash and prefix-cache extra...
Collapse the non sorted axes to pick the contiguous sort kernel ml-explore/mlx
mx.sort and mx.argsort returned silently wrong results on transposed views on GPU, with no error and no crash, because non-sorted axes were not collapsed onto the contiguous kernel. The reproducer (a 3x4x8 array through mx.swapaxes then sort along axis=-1) compared False before and True after. If you sort or argsort non-contiguous views, the previous results you based anything on are not trustw...
llama: fix clef head reads past 2GiB on windows (#18777) ollama/ollama
The Clef head read its weights with std::ifstream::seekg, and on the Windows llama-server build (libc++ 18.1.8) that seek truncates offsets to 32 bits, so clef-flash's clef.* tensors at about 9.5GB loaded backbone bytes as head weights and every /v1/systemone request failed with "Clef: non-finite logit". clef:27b was unaffected because its head and output.weight stay under 2GB; the read now goe...
model: add multimodal embeddings ollama/ollama
Ollama implemented the EmbeddingGemma2Model architecture on the MLX runner: a 24-layer bidirectional text encoder with PLE, shared gemma4 vision and audio towers, and mean-pool plus L2 output. /api/embed now accepts per-item media through input dicts, so one embed call can mix text and image or audio items rather than requiring separate requests.
Fix float16 overflow in clip_grad_norm ml-explore/mlx
clip_grad_norm squared and summed gradients in their own dtype, so any fp16 gradient above roughly 256 overflowed to inf, making the total norm inf, the normalizer zero, and every gradient silently zero; the reporter confirmed an fp16 value of 400 returning inf on the released 0.32.3 core. The fix sums fp16 squares in fp32 and scales in fp32 before casting back, keeping clipped gradients in the...
RPC: add `-sm tensor` ggml-org/llama.cpp
RPC gains -sm tensor, tested across 2x Sparks connected over RDMA, which lets tensor parallelism span machines instead of confining each GPU to its own layer slice. It needs async graph_compute, a custom all_reduce, a graph uid cache like CUDA's, and set_tensor_2d/get_tensor_2d in the RPC layer, so it is a reviewable proposal more than a drop-in switch.
[DeepSeek-V4.1] Unify mHC into one state machine and drop medium-batch fusion variants sgl-project/sglang
DeepSeek-V4.1's manifold-constrained hyper-connections now live in one module, deepseek_v4_mhc.py, with a single state machine for cross-layer state and one dispatch point per mHC step, replacing contextvar hooks threaded through unrelated code and dropping the medium-batch fusion variants. The PR is marked as generated by Claude, which is worth knowing before you read the design notes.
Action items
- → Patch or firewall exposed vLLM multimodal endpoints: request-controlled mm_processor_kwargs.code_revision reaches cod... vllm-project/vllm [immediate]
- → Re-run any results derived from mx.sort or mx.argsort on non-contiguous or transposed views on current MLX, since ear... ml-explore/mlx [immediate]
- → Move Windows llama.cpp clef-flash deployments off builds that read the Clef head past 2GiB ggml-org/llama.cpp [plan]