Fetching from the wire…
OSS2026-08-28 · source-backed
The day's highest-scoring r/LocalLLaMA post points out the deal takes the llama.cpp and ggml copyright along with the team Hugging Face hired in February 2026, including Georgi Gerganov (r/LocalLLaMA). The top reply at 957 upvotes is "If it happens, we shall fork and move on. It is the way of things," and a separate thread argues Nvidia has an incentive to drop support for older cards like the V100 that llama.cpp keeps alive. The counter-argument: Nvidia and llama.cpp maintainers have been collaborating on multi-GPU and tensor-parallel work in ggml for months.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
Announced July 27 with Microsoft, IBM, Red Hat, Palantir, CrowdStrike, Cloudflare, Databricks, Hugging Face, LangChain, Nous Research, Reflection AI, Thinking Machines Lab, SpaceXAI and the Linux Foundation. Huang's framing is pointed: during the Hugging Face incident "closed...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
Poolside told investors the license is non-exclusive, covering the system used to train its Laguna open model, plus job offers to 109 employees and a separate $1B investment at a $12B pre-money valuation. The letter insists this is neither acquisition nor acquihire, and the th...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.