vLLM's request-controlled caches, and Vulkan sparse attention for quantized K/V
- vulkan: sparse flash attention for quantized K/V ggml-org/llama.cpp
$ git log --author="fxgsell" --all --oneline
Every morning, RepoJournal's newsroom reads what shipped across the open-source projects developers depend on — and writes the wire. Your commits kept making the news. This is your clippings file: the stories where your work was the story.
1
story filed
1
daily wire
$ grep -rn "author: fxgsell" wires/ | sort -r