$ the-wire · showcase
llama.cpp sanitizes invalid UTF-8 in the parser AST
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
Parser robustness, CUDA attention tuning, and memory bounds dominate a day of work aimed at the long tail of production inference: malformed model output, sparse attention on Qwen4, and an indexer that OOMs with room to spare in the cache.
common/peg : handle invalid utf-8 sequences in the AST ggml-org/llama.cpp
Models occasionally emit invalid UTF-8, which used to fail during parsing; the parser now records those runs and exposes sanitized_text() on every AST node, replacing invalid sequences with the maximal subpart per Unicode recommendations. Tool names and args are deliberately left un-sanitized on the assumption that grammar constraints keep them valid, so anything feeding tool call text through ...
CUDA: enable sparse fa for qwen4 ggml-org/llama.cpp
Sparse FlashAttention is now enabled for Qwen4 in the CUDA path, switching on once context reaches 32768 by taking the union of tokens in use at ncols1=8. The author notes the model's attention is otherwise unoptimized, with the entire kv-cache re-scored on every step, and flags that as still open.
[DeepSeek-V4.1] Bound dense prefill indexer memory sgl-project/sglang
The DeepSeek-V4.1 dense prefill indexer could OOM even with KV cache headroom, because DeepGEMM returns FP32 scores shaped [query tokens, visible context entries] that stay alive while Boolean candidate masks are built in pieces and concatenated. The fix bounds that memory; FP4 inputs are unaffected in precision.
Use pinned memory for asynchronous sampling metadata transfers sgl-project/sglang
Sampling metadata now moves over pinned host memory: stage penalty parameters, sparse logit bias, custom-processor masks, and raw token-ID logprob indices use nonblocking device copies, and stop-token IDs are padded on CPU and transferred once per batch. Sampling semantics, dtypes, device placement, and platform-specific pinning support are preserved, so this is a throughput change rather than ...
metal : support arbitrary hc in dsv4_hc_pre ggml-org/llama.cpp
Metal's dsv4_hc_pre kernels now accept arbitrary hc values instead of the hardcoded hc = 4. Elsewhere in the day's long tail: CUDA FlashAttention tuning for Gemma 4 head sizes 256 and 512, CUDA graph capture for the DeepSeek-V4.1 vision encoder, Humming quantization gaining Hadamard transforms and SM120/SM121 support, a dead-kernel cleanup, an NVFP4 MoE dispatch fix where create_weights and app...