Fetching from the wire…
Public story · 2026-03-07 · source-backed
Nathan Lambert's deep dive on OLMo Hybrid makes a compelling case that the transformer-only era may be ending. The paper formally proves hybrid models solve problems (related to code evaluation) that neither transformers nor GDN can solve alone. Combined with Qwen 3.5 and Kimi Linear taking similar approaches, Lambert sees a "resurgence of truly open models" driving architecture innovation that closed labs can't match. Critical caveat: inference tooling is 3-6 months behind, so production deployment requires workarounds today. (Interconnects)
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
AI2's fully open 7B model replaces 75% of attention layers with Gated DeltaNet linear recurrence, matching OLMo 3 accuracy on MMLU with 49% fewer tokens. Nathan Lambert's analysis highlights the theoretical contribution: hybrid models formally solve problems that neither pure...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
1. Set Up Cursor Automations (intermediate) — Event-driven agents from PagerDuty/GitHub/Slack triggers with isolated sandboxes. Cursor Blog 2. Apply Context Engineering to Cut Agent Costs 60-80% (advanced) — Hierarchical token budgets, dynamic tool filtering (max 15), automati...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.