Local LLMs
QWEN3.5 SPECULATIVE DECODING HITS 70% SPEEDUP ON LLAMA.CPP, OLLAMA ROUTES CLOUD MODELS
Qwen3.5's MTP head is now live across both Ollama and llama.cpp, delivering 24-70% inference speedup on dense models while Ollama adds smart cloud-model routing to end user confusion.
read --wire →