$ the-wire · showcase
OpenVINO 2026.4.1 lands in llama.cpp, vLLM frees CUDA graph memory on sleep
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
Three inference stacks shipped headroom today: OpenVINO's 2026.4.1 bump brings Gemma and Qwen MoE performance work to llama.cpp, a new vLLM flag lets sleep mode actually reclaim CUDA graph memory on MoE deployments, and sglang untangles mamba caching with the radix cache disabled.
ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. ggml-org/llama.cpp
The OpenVINO backend moves to 2026.4.1 with tighter MoE expert fusion for Gemma and Qwen, expanded op coverage including common MTMD ops, and a clarified --list-devices plus GGML_OPENVINO_DEVICE selection path. Also along for the ride: GPU fixes, build warning cleanup, and stale OpenCL check removal, so Intel GPU users get both speed and fewer surprises at device pick time.
[Feature] Release the CUDA graph pool on sleep vllm-project/vllm
Sleep mode used to unmap weights and KV cache while CUDA graphs kept their private tensor pools pinned, which on large MoE deployments meant GiBs per GPU that sleep could not release. The new opt-in sleep_mode_offload_cudagraph, off by default, captures graphs into a dedicated cuMem-backed pool so sleep frees the physical memory and wake remaps the same virtual addresses without recapturing; wi...
[mem_cache] Run mamba models on `UnifiedRadixCache` when the radix cache is disabled sgl-project/sglang
With --disable-radix-cache, mamba models (GDN, Mamba2, KDA, lightning, plus MiniCPM SALA and the PD decode side) now run on UnifiedRadixCache's disabled mode instead of ChunkCache, and the mamba component owns mamba-state release in both modes. The page_size == 1 requirement now applies only when the tree actually caches states, which also gets --disable-radix-cache --enable-streaming-session s...
[Bugfix][Responses API] Reuse streamed item ids in final harmony response vllm-project/vllm
Streamed /v1/responses rebuilds response.completed from harmony messages, and that rebuild was minting fresh item ids and function call_ids that no longer matched the ones already sent in response.output_item.added/done, so strict clients aborted the stream. The fix records the items sent for each completed harmony message and reuses their ids; the non-harmony path was addressed separately.
[sgl-router] Repair KV-event sequence gaps from the engine's replay socket sgl-project/sglang
ZMQ drops KV events at the publisher's high-water mark, and a lost BlockRemoved previously left the router crediting a worker with blocks it no longer held until the next AllBlocksCleared, with sgl_router_kv_event_batches_lost_total as the only trace. Per-(worker, dp_rank) subscribers now track their own last seq and, when /server_info advertises replay_endpoint_port_base, pull the missing even...
Unroll TK=4 keys per iter in sdpa_vector_2pass_1 (#4596) ml-explore/mlx
The rest of the day is maintenance: MLX unrolls the generic sdpa_vector_2pass_1 key loop four keys per iteration (the GQA specialization is untouched), llama.cpp adds common_is_tty() and quiets deprecated warnings on Windows, vendors cpp-httplib 0.59.0, and vLLM bumps the Transformers version used in CI.