$ the-wire · showcase
Ollama eases off typical_p, llama.cpp lands Vulkan int8 tensor cores
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
Ollama downgraded typical_p from a hard API rejection to a warning, while llama.cpp put integer tensor cores to work on AMD and Qualcomm Vulkan paths and vLLM began reporting reasoning tokens it had been silently zeroing.
api: deprecate typical_p (#18627) ollama/ollama
Ollama's API used to reject requests carrying the deprecated typical_p parameter; it now logs a warning instead, so clients that still send it keep working. Model creation still rejects the parameter, meaning the deprecation is only half-parsed by the API surface.
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (#27952) ggml-org/llama.cpp
llama.cpp added an int8 cooperative-matrix quantized matmul shader for AMD RDNA3 and RDNA4, with q8_0 support, preloaded scales, and workgroup scheduling for cache. Most devices' int8 tensor cores have higher throughput than the fp16 instructions the Vulkan backend previously used for all matrix-core GEMM.
[Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs vllm-project/vllm
The per-sequence Python loop that saved block-aligned SSM states in MambaMixer2.conv_ssm_forward is gone, and with it the four .tolist() calls that forced GPU-to-CPU syncs on every prefill. The writes now collapse into one pair of index tensors and a single indexed copy.
[Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers vllm-project/vllm
The Wire's own reporting caught this one: /v1/chat/completions always reported completion_tokens_details.reasoning_tokens as 0 for gpt-oss, because HarmonyParser never overrode count_reasoning_tokens and the call fell through to GptOssReasoningParser. /v1/responses had been counting correctly, so this reconciles the two endpoints for Harmony, DeepSeek-V3, and Step3.
Fix mixed chunk prefill with DP speculative coordination sgl-project/sglang
SGLang removed a guard that raised "Local mixed prefill/verify is not implemented" when a prefill rank held local decode requests, so mixed-chunk ranks can now run prefill with one-token local decode progress while decode-only peers keep speculative decoding. The unit test asserting the old rejection was deleted and replaced with a standalone GPU regression test. Elsewhere: MLX fixed CPU/GPU so...