Fetching from the wire…
Models2026-03-23 · source-backed
MiMo-V2-Pro (released March 18) scores 61.5 on ClawEval — Claude Opus 4.6 scores 66.3, GPT-5.2 scores 50.0 — with over 1T total parameters (42B active), a 1M-token context window, and free availability on OpenRouter. It ran as "Hunter Alpha" in stealth on OpenRouter, processing over 1T tokens before identification. The companion Flash model (309B, open source) beats most models at its weight class. A phone company is now third in the world on agent benchmarks.
Each link below shares sources, entities, or timing with this story.
Microsoft announced Critique on March 30. Here's how it works: when you use M365 Copilot Researcher, GPT drafts the initial research response. Then Claude reviews it for accuracy, completeness, and citation quality. You only see the final result after both models have had thei...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
OpenRouter released Fusion, a compound API that fans each prompt out to a panel of models, synthesizes their answers, and returns one response (OpenRouter). On Perplexity's DRACO deep-research benchmark, 100 tasks across 10 domains, a Fable 5 + GPT-5.5 fusion scored 69.0% vers...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.