$ the-wire · showcase
Ollama drops the CLI account step, MLX fixes foreign buffer ownership
By RepoJournal · Filed · About Local LLMs · Composed from the cited sources · methodology
A day of correctness work across the stack: Ollama shortens onboarding and repairs Windows manifests, MLX fixes buffer ownership and an FFT shape bug, and llama.cpp makes temperature-zero sampling actually greedy.
manifest: avoid symlinks on Windows ollama/ollama
Named manifests are now stored as regular files on Windows, where symlinks create successfully but fail when opened; existing symlinks are repaired lazily on read using the verified blob and the existing copy path. Locking stays consistent across platforms and pruning no longer reacquires the store lock, which fixes the model-load failures tracked in #18847.
Fix foreign buffer ownership ml-explore/mlx
The buffer and its deleter are now coupled in the data, and donation no longer fires when MLX does not own the underlying buffer: a writable numpy.ndarray is no longer overwritten by calls like mx.exp(np_array), and the same applies to Python bytes, mmap arrays, and DLPack buffers. The first op in those chains now allocates instead of writing into memory MLX does not own.
sampling : use greedy selection for eligible temperature-zero chains ggml-org/llama.cpp
Eligible temperature-zero backend sampler chains still ran distribution sampling even though token selection is deterministic; they now use greedy selection while retaining preceding filters, with the selected index mapped back to its vocabulary token ID. All other sampling configurations keep the existing path. The PR reports SPEED-Bench coding results, 80 conversations and 89 turns, against s...
[Bugfix][DSv4.1] Skip SWA bounded replay for KV loads that carry the window vllm-project/vllm
SWA bounded replay was replaying the last window of every prefix hit, including KV delivered by P/D connectors that already carry the request's own sliding-window KV, so DeepSeek-V4.1 decode recomputed the last 128 prompt tokens per request. The whole-block rounding applied to external hits also shrank decode's allocation below what prefill exports, breaking NIXL's end-to-end transfer.
model : add LiquidAI/d1-omni-600M decision model ggml-org/llama.cpp
llama.cpp gains LiquidAI/d1-omni-600M, an omni decision model taking audio, images, and text, with GGUFs published for testing. It depends on the landed work in the referenced ggml-org/llama.cpp PR. Elsewhere the long tail is cleanup: hexagon GELU/GELU_ERF accuracy and alloc_buffer_n, a MLX_CONV_WINOGRAD=0 switch for conv2d, HiCache's SeaweedFS L3 backend, and duplicate block-size defines dropp...