$ the-wire · showcase
Ollama gates System One on declared capabilities, llama.cpp fixes a stale-UI trap
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
Ollama moves capability decisions into the Modelfile while llama.cpp, vLLM, sglang, and MLX ship correctness and performance work underneath: a prefix-cache collision fix, a service worker that refuses to die, and FP32 Hexagon kernels.
create: support explicit model capabilities ollama/ollama
Modelfiles gain CAPABILITY declarations and create requests an additive capabilities field, preserved across GGUF and safetensors creation, inheritance, and Modelfile export. System One requests are now scheduled only when the model declares the decision capability, replacing the old Qwen architecture and renderer metadata match; GGUF-only scoring stays in place until the separate MLX runtime w...
[Core][BugFix] Tag prefix-cache extra keys by source vllm-project/vllm
Prefix-cache block hashes previously mixed LoRA names, cache_salt, multimodal identifiers, and prompt-embeds digests into one untagged extra_keys tuple, so a LoRA request for adapter foo and a base-model request with cache_salt="foo" produced the same first-block hash for the same tokens. A request could then be served KV computed with another request's weights, and a caller could pick a salt l...
server : remove the built-in UI's service worker when the UI is not served (#29565) ggml-org/llama.cpp
With --path or --no-ui, /sw.js returned 404, and a 404 does not unregister a service worker, so browsers kept serving the cached built-in UI. The server now returns a worker that unregisters itself, clears its caches, and reloads open tabs, while a sw.js in the --path folder still takes precedence.
[Bugfix][Responses API] Build streamed final response from streamed items vllm-project/vllm
In non-harmony streaming, /v1/responses rebuilt response.completed by re-parsing the full output instead of reusing what it streamed, so the two disagreed: items got new ids, a function call's call_id changed, and clients continuing with previous_response_id referenced a call_id the stored response did not have. Streamed items could also vanish (a whitespace-only message before a tool call) and...
[Security] Bump nltk, aiohttp, pillow, and datamodel-code-generator (#59249) vllm-project/vllm
Upstream pins were raised for nltk, aiohttp, pillow, and datamodel-code-generator, so images and environments building vLLM from source should pick up the corrected pins rather than resolving the old ones. Those four are the security bump to take before the next build.
[Kernel][Perf] Add TP=2/4/8 per-rank shapes to the sm_120 batch-invariant matmul table vllm-project/vllm
The sm120 batch-invariant matmul table only carried TP=1 shapes for Qwen3-1.7B/4B/8B, so VLLM_BATCH_INVARIANT=1 with tensor_parallel_size > 1 left every linear layer on RTX PRO 6000 and RTX 50 on the default 128x128x64, 8 warps, 3 stages config, leaving the roughly 2.9x batch-invariant overhead from the TP=2 rows unchanged. Per-rank TP=2/4/8 shapes are now in the table.
[Kernel] Expose FlashMLA kv_format in sgl-kernel sparse decode sgl-project/sglang
FlashMLA's V3.2-no-RoPE and V4.1 KV caches are both 528 bytes per token at head_dim 512, so shape alone cannot separate them, and sgl-kernel always passed no format, making FlashMLA read a V4.1 cache as V3.2-no-RoPE. sparse_decode_fwd and flash_mla_with_kvcache now accept an optional kv_format.
docs: document System One API ollama/ollama
A Decision guide and System One API reference now cover choice, yes/no, and scoring questions, including local availability, request and response fields, and limits.
Action items
- → Rebuild vLLM images and environments to pick up the nltk, aiohttp, pillow, and datamodel-code-generator security pins vllm-project/vllm [immediate]