Fetching from the wire…
Public story · 2026-02-12 · source-backed
Opus 4.6 (Thinking) leads SWE-bench Verified at 79.2%, but SWE-Bench Pro (~23%) and SWE-EVO (19-21%) reveal significant gaps. Current coding agents excel at well-defined single-repo issues but struggle with cross-language, cross-repo, evolutionary tasks — the reality of production engineering.
Each link below shares sources, entities, or timing with this story.
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the buil...
xAI launched Grok 4.5 and Grok Build on July 8, trained partly on Cursor developer-session data. The numbers are loud: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, priced at $2/$6 per million tokens. On a single coding task that works out to roughly $2.49 versus $11.80...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Anthropic launched Claude Sonnet 4.6 claiming performance comparable to Opus 4.5 at $3/$15 per million tokens (vs. Opus's $5/$25). SWE-bench Verified: 79.6% (near Opus 4.6's 80.8%). OSWorld-Verified: 72.5% (tied with Opus 4.6's 72.7%). 1M token context window in beta. Now the...
GLM-5.1 scored 58.4% on SWE-Bench Pro. Opus 4.6 scored 57.3%. GPT-5.4 scored 57.7%. Read those numbers again. An open-weight, MIT-licensed model now leads the most rigorous coding benchmark we have. This isn't a narrow win on a cherry-picked eval. SWE-Bench Pro tests real-worl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.