Fetching from the wire…
Public story · 2026-03-20 · source-backed
A high-engagement thread (143 comments) challenges the agent-and-coding obsession: the original use case that drew many practitioners was superior knowledge retrieval over search engine noise, largely unsolved three years later. Models optimized for agentic coding often sacrifice the contextual knowledge depth that makes LLMs useful for research. A real gap in the development trajectory.
Each link below shares sources, entities, or timing with this story.
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
r/LocalLLaMA benchmarking (169 upvotes, 92 comments) documents Qwen 3.5 397B outperforming gpt-oss 120b, StepFun 3.5, MiniMax M2.5, and Super Nemotron 120B on coding tasks. Relevant for heterogeneous model routing in agentic pipelines — 397B may justify the inference cost for...
The post argues Ollama went over a year without crediting llama.cpp in its README while a license-compliance issue sat 400+ days without a maintainer response, that llama.cpp runs 1.8x faster (161 against 89 tokens/second) with 30-50% CPU gaps, and that the mid-2025 move to a...
A 1,134-upvote r/LocalLLaMA post pushes back on the claim that n-gram tables let you run 1T+ models with 980B parameters offloaded to SSD (r/LocalLLaMA). An engram is an embedding table keyed on the last two or three tokens rather than one token ID, so "New York" gets a memori...
An r/LocalLLaMA thread (91 upvotes, 74 comments) documents the shift. Pi's system prompt is under 1,000 tokens vs OpenCode's 10K+, with faster startup and better local model performance on Mac with MLX. Reveals a practitioner split between "everything-connected" and "fast-and-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.