Fetching from the wire…
Infra2026-06-13 · source-backed
PR #21038 adds graph-level Hadamard rotation of Q/K/V before caching, doing attention in the rotated space then rotating back, which makes standard quant types far more accurate in the KV cache. The benchmarks are not subtle: Qwen3 0.6B q5_1 KV perplexity drops from ~61.7 to ~14.1, and q4_1 from ~212 to ~22.3. For self-hosters this is a near-free path to longer contexts in less VRAM, and r/LocalLLaMA is tracking it as the successor to earlier TurboQuant work.
Each link below shares sources, entities, or timing with this story.
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Anthropic, OpenAI, Google, Meta, Microsoft and Mistral are all Section 1 signatories of the EU Code of Practice on Transparency of AI-Generated Content, and the 315-upvote, 246-comment thread centers on whether open-weight models from those companies carry watermarking too (Eu...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.