$ the-wire · showcase
llama.cpp v0.5.0 ships, vLLM discloses five engine-fatal request-handling flaws
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
llama.cpp cut v0.5.0 with backend and server work, while vLLM's Rust frontend and multimodal stack got hit with a set of unauthenticated denial-of-service and integrity disclosures that server operators cannot ignore.
v0.5.0 ggml-org/llama.cpp
by github-actions[bot]
The release rolls up backend correctness and performance work with broader model coverage, including HRM-Text (DFM Mimir 1B) support and MiMo-V2.6 and HunyuanOCR conversion, plus multi-address HTTP binding and image outputs from function calls. CUDA conv2d now uses an implicit GEMM path, and the release notes flag Metal and chat parser fixes aboard.
Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites) vllm-project/vllm
Errors raised by structured-output requests escape the request boundary and terminate the shared EngineCore, taking down every tenant on the engine; the disclosure lists three affected sites. Treat this as a multi-tenant availability risk rather than a per-request one.
Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens in the vLLM Rust frontend metrics middleware (unauthenticated denial of service) vllm-project/vllm
The Rust frontend's metrics middleware interpolates attacker-controlled HTTP method tokens straight into Prometheus label values, so unauthenticated traffic grows label cardinality without bound. Cardinality growth like that degrades or kills the metrics pipeline serving the whole server.
Scale-out disaggregated multimodal transport trusts caller-supplied features — shared EngineCore denial of service, encoder-cache poisoning, and transport integrity loss (5 sites) vllm-project/vllm
Scale-out disaggregated multimodal transport trusts caller-supplied features across five sites, which the disclosure says enables shared EngineCore denial of service, encoder-cache poisoning, and loss of transport integrity. Anything routing multimodal requests across replicas is exposed.
Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the num_frames ceiling does not reach vllm-project/vllm
The Qwen2-VL and Qwen3-VL video samplers bound on a request-controlled max_frames, which the num_frames ceiling never reaches, so callers can drive sampling past the intended limit. Story 13 describes the same shape in GLMGA video sampling: request-driven CPU and memory exhaustion.
metal : key the fa-vec tuned table by family instead of SKU ggml-org/llama.cpp
In the long tail: the Metal fa-vec tuned table got rekeyed from device_id to gpu_family, collapsing 3884 rows across 19 SKUs to 1033 rows across 4 Apple GPU families, and sglang's Wan2.2 DiT FP8 attention path replaced its multi-kernel per-tensor quant reference with the AITER backend.
Action items
- → Apply the vLLM fixes for the shared EngineCore and Rust frontend metrics denial-of-service disclosures before exposin... vllm-project/vllm [immediate]
- → Upgrade vLLM to a build that bounds Qwen2-VL/Qwen3-VL video sampling by num_frames rather than the request-controlled... vllm-project/vllm [plan]