$ the-wire · showcase
Ollama folds structured output into one pass, llama.cpp binds multiple addresses
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
Ollama cut the double-prefill cost of JSON-constrained thinking models and sped up Qwen prompt processing on MLX, while llama.cpp made llama-server bindable to several addresses at once and shaved the iq1_m quantizer from O(B³) to O(B²).
server: apply structured outputs in a single pass on thinking models ollama/ollama
A format on a thinking model used to run two generations: an unconstrained one the server cancelled once the parser reported content, then a re-rendered prompt under the grammar. Parsers now report the strings that end a response, so the same request completes in a single pass, which removes the restart's second prefill, the chunk dropped at the boundary, and a stray first token that could leak...
mlx: speed up Qwen 3.8 prompt processing ollama/ollama
dhiltgen's MLX path uses the gated-delta kernel for long scans and folds dense MLP global scales into SwiGLU, lifting prompt processing on an M5 Max from 715.2 to 848.5 tokens/s at 2k context (+18.7%), 695.6 to 828.1 at 8k, and 704.6 to 802.6 at 16k.
server: Add support for binding to multiple addresses ggml-org/llama.cpp
Binding llama-server to a VPN address previously made it unreachable through localhost, breaking clients that default there. --host now takes a comma-separated list of IP addresses and Unix socket paths, one listener per address sharing the same HTTP worker pool, so a single server answers on both; with --port 0 the first TCP listener selects the port.
ggml-quant: Quant creation speed, IQ1_M build prefix sums once per block ggml-org/llama.cpp
iq1_m recomputed weighted sums across all 16 elements for every candidate; the quantizer now builds prefix sums once per block and scores in constant time, dropping runtime from O(B³) to O(B²). The iq1_s quantizer already did this, and float sum ordering means the output is not byte-identical.
mlx: select Gemma 4 image resolution dynamically ollama/ollama
The MLX Gemma 4 path replaces the fixed checkpoint image budget with per-image selection across the 70, 140, 280, 560, and 1120 budgets, picking the publisher resize grid closest to the input resolution with aspect ratio accounted for. High-resolution documents keep more detail and smaller images spend fewer tokens, with no new API parameter.
[dsv4.1]Optimize FP4 indexer by skipping invisible tiles sgl-project/sglang
sglang's legacy FP4 indexer replayed a capacity-sized CUDA Graph grid even on short contexts: in a measured shape with 1,048,896 output columns and 8,704 visible positions, over 99% of columns were invisible yet still ran dequantization, dot products, weighting, and reduction. The indexer now skips arithmetic for fully invisible tiles and writes -inf directly to valid output addresses.
[HiCache] Remove the unused HiRadixCache sgl-project/sglang
Elsewhere on the long tail, sglang deleted the unused HiRadixCache along with its helpers and tests, since HiCache has gone through UnifiedRadixCache since v0.5.19, and dropped the deprecated SGLANG_ENABLE_UNIFIED_RADIX_TREE env now that the unified radix tree is the default; vllm removed its AllSpark INT8 W8A16 GEMM backend and dropped the model_runner_v2 test from Intel GPU CI, and llama.cpp ...