$ the-wire · showcase
vLLM locks down multimodal kwargs, llama.cpp cleans up failed restores
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
vLLM now rejects untrusted per-request multimodal processor overrides by default, while llama.cpp closes a state-corruption path after failed sequence restores and MLX trims attention memory on NAX GPUs.
[Security] Gate per-request multimodal processor kwargs vllm-project/vllm
API clients could previously send per-request mm_processor_kwargs and media_io_kwargs that flowed straight into multimodal media loading and processor construction, letting untrusted overrides force extreme image or video settings and exhaust memory. Non-empty per-request kwargs are now rejected by default, with a new --trust-request-mm-kwargs flag on BaseFrontendArgs for deployments that need ...
llama : fix K/V and recurrent state cleanup after failed restores ggml-org/llama.cpp
A failed per-sequence state restore could delete sequence metadata while leaving K/V or recurrent-state tensor data written by that restore, and buffer-backed restores could leave deferred writes pending to be applied after cleanup. The fix clears affected tensors and discards deferred writes before the reader is destroyed, so a failed restore no longer leaves stale state behind to corrupt late...
Improve metal memory usage for SDPA D512 ml-explore/mlx
Attention at D512 on NAX GPUs previously held a larger working set than necessary; the branch cuts peak memory by 1.91 GB (7.99%) on mlx-community/gemma-4-31b-it-nvfp4 at 16K and by 4.22 GB (30.6%) on the 12B model at 65K in the PR's measurements. Larger context windows on Gemma 4-class models get headroom without changing results.
[Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights vllm-project/vllm
RoutedExperts.load_weights scanned the entire expert mapping for every checkpoint tensor, making matching quadratic in expert count; profiling of Qwen3.8-Flash-Next-NVFP4 on a DGX Spark attributed roughly 25 seconds of weight loading to that loop. A new _index_expert_mapping helper groups entries by the checkpoint's literal logical expert ID, so load time stops scaling with expert count squared.
cuda: add F16 input to the FWHT (#29096) ggml-org/llama.cpp
The CUDA FWHT accepted F32 input only, forcing a converted copy; the source type is now a template parameter so the kernel reads F16 directly, with supports_op accepting an F16 src1 against an F32 src0 for the Hadamard hint and every other F16 src1 against a non-F16 src0 still refused. The F32 path is unchanged, and ggml_cuda_op_mul_mat_use_fwht remains the single predicate shared by supports_o...
Action items
- → Set --trust-request-mm-kwargs on vLLM servers that rely on client-supplied mm_processor_kwargs or media_io_kwargs; th... vllm-project/vllm [immediate]
- → Update to a llama.cpp build containing the failed-restore state cleanup if you serve multiple sequences per context ggml-org/llama.cpp [plan]