$ the-wire · showcase
vLLM's request-controlled caches, and Vulkan sparse attention for quantized K/V
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
Two remote-triggerable resource bugs land in vLLM's queue while llama.cpp turns on sparse flash attention for quantized KV caches, the fix that matters most to anyone running long-context inference on non-NVIDIA GPUs.
MOSS-Audio request-controlled processor cache permits remote memory exhaustion vllm-project/vllm
A request-controlled processor cache in MOSS-Audio lets a remote caller exhaust memory, so a single crafted request is enough to take a serving instance down. Treat this as the day's most urgent item if you expose MOSS-Audio endpoints to untrusted traffic.
An oversized min_tokens in one request hangs the vLLM engine while /health stays green vllm-project/vllm
An oversized min_tokens in one request hangs the engine while /health keeps reporting green, which means your orchestrator will not restart the pod for you. Until a fix lands, cap min_tokens at the gateway rather than trusting readiness probes.
vulkan: sparse flash attention for quantized K/V ggml-org/llama.cpp
Vulkan's sparse flash attention previously activated only when K and V were f16, so quantized-cache models like Qwen3-Flash-Next QSA and DeepSeek-style sparse attention ran dense attention over the entire context despite keeping only about 2k cells. The scalar and coopmat1 paths already dequantize each gathered row through the sparse index, so the gate just needed to admit quantized K/V; coopma...
mlx: mitigate high latency after GPU idle ollama/ollama
ollama now carries the MLX residency-refresh patch and enables it every second in the runner, preserving explicit environment overrides, which targets the high-latency stall after the GPU has sat idle. The same author also landed model-lookup and MLX decision-request overhead cuts, avoiding decodes of unrelated manifests and reusing small Metal scratch buffers between decision requests.
[HiCache] Route --file-storage-path to the file storage backend sgl-project/sglang
--file-storage-path is parsed into server_args.file_storage_path but nothing reads it: the file HiCache backend honors only SGLANG_HICACHE_FILE_BACKEND_STORAGE_DIR and otherwise writes to /tmp/hicache, so L3 silently lands on the wrong disk or on a tmpfs that is far smaller than intended.
Fix metal validation error when passing empty vector ml-explore/mlx
Elsewhere: MLX replaced Black with ruff, valid scalar and zero-index GPU ops got typed placeholder bindings to stop Metal validation aborting Gather::eval_gpu, and default.metallib lookup now resolves under swift test with the Swift Build system, a case where MLX previously killed the entire test process on the first GPU call.
Action items
- → Restrict or firewall MOSS-Audio endpoints until the request-controlled processor cache is capped vllm-project/vllm [immediate]
- → Cap min_tokens at the gateway before relying on /health, which stays green while the engine hangs vllm-project/vllm [immediate]