Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Compiled 2026-09-09 · source-backed
Ollama runs 1.8x slower than llama.cpp at 89 vs 161 tokens/second.
Source findingllama.cpp enables speculative-decoding state rollback for Kimi-K3
Source findingllama.cpp Vulkan backend fuses DeepSeek-V4 ops reducing decode time by 32%
Source findingllama.cpp added HF-to-GGUF conversion support for Qwen3.5.
Source findingllama.cpp added HF-to-GGUF conversion support for Qwen3-Next.
Source findingNInfer and llama.cpp compared on concurrent serving; llama.cpp capped at parallel=1
Source findingA llama.cpp fork added support for Q2_B3 base-3 ternary weight packing format.
Source findingllama.cpp 0.4.0 adds initial support for Qwen3.8-Flash-Next.
Source findingllama.cpp 0.4.0 runs on ggml 0.23.0 with lazy tensor reading and quantizer RAM capping.
Source findingllama.cpp 0.4.0 adds support for DeepSeek-V4 with sparse flash attention.
Source findingllama.cpp achieves 5x speedup on Vulkan IQ3_S at batch 8
Source findingllama.cpp fuses QKV and FFN matmuls on Qualcomm Hexagon HMX
Source finding