$ the-wire · showcase
llama.cpp spec sampling goes probabilistic, vLLM trims GLM-5.3 attention overhead
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
llama.cpp reworks speculative decoding to stop discarding the drafter's distribution, while vLLM and sglang land targeted kernel and cache fixes that change what existing deployments do at temperature above zero and after prefix eviction.
Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP ggml-org/llama.cpp
Model-based drafters used to run the sampler, throw the result away, and hand the target a single top candidate, which the target then had to draw again at temperature above zero; the drafter is now probabilistic and the target verifies by rejection sampling, with a flag to enable probabilistic draft sampling (default stays greedy) and argmax fallback for grammar-constrained requests. The PR st...
[GLM5.3 Perf] Reuse sparse MLA index conversion across layers, 3.5~3.9x kernel performance improvement vllm-project/vllm
GLM-5.3 shares top-k indices across layer groups, but FA3 sparse MLA was re-deriving the same physical KV slots per layer; converting once per group cuts conversions from 78 to 21 per main-model forward, cited as a 3.5 to 3.9x kernel improvement.
fix: preserve QSA indexer state through HiCache sgl-project/sglang
QSA compressed index keys live outside full_kv_pool, so the hybrid HiCache stack was restoring KV and Mamba state without the indexer state that selects attention blocks. After device eviction that let a prefix hit read valid KV through stale QSA block selections and produce incorrect output, a correctness fix for anyone running prefix caching on the hybrid stack. The related PLE slot-state fix...
CUDA: fuse shared experts into MMVQ ggml-org/llama.cpp
Shared experts are now fused into the MMVQ path; on Qwen3.5-35B-A3B-Q8_0 on RTX Pro 6000 the gain runs from +0.42% at microbatch 4 to +5.06% at microbatch 5, with +4.54% at microbatch 2.
Size the VMM graph-input exchange by the widest input across ranks sgl-project/sglang
The VMM graph-input exchange packed expandable-segment chunk indices into a fixed-width struct sized for 16 chunks, so a rank whose input spanned more raised "Too many VMM chunks for graph input" while its peers blocked in the all-gather. The exchange is now sized by the widest input across ranks, which removes a hang that depended on allocator segment size rather than the model.