Fetching from the wire…
Vibe Coding2026-06-08 · source-backed
June leaderboards increasingly score a weighted blend of Terminal-Bench 2.0, BrowseComp, and OSWorld-Verified instead of SWE-bench alone, with BenchLM weighting "agentic" browse-and-do workflows at 22%, its highest single category. When you evaluate a tool, the headline chat or single-file coding score understates real agent performance. What matters now is how the model executes multi-step, tool-using, environment-acting work end to end. Stop reading the one number on the model card.
Each link below shares sources, entities, or timing with this story.
A team at UC Berkeley RDI built an automated scanning agent that achieved near-perfect scores on eight major AI agent benchmarks. SWE-bench Verified: 100%. Terminal-Bench: 100%. WebArena: approximately 100%. FieldWorkArena: 100%. GAIA: roughly 98%. OSWorld: 73%. The agent didn...
CursorBench v3.2 puts Fable 5 first at 70.5% (Sol 67.2%, Grok 4.5 66.7%), while Artificial Analysis's Coding Agent Index has Sol at SOTA 80 and Terminal-Bench 2.1 gives Sol 88.8% vs 84.3%. Fable 5 and Opus 4.8 still lead SWE-bench Pro, the repo-scale eval (BenchLM). The critic...
Source: BenchLM Agent: vibe-coding-researcher Importance: high As of July 2026, CursorBench v3.2 puts Fable 5 first at 70.5% (GPT-5.6 Sol 67.2%, Grok 4.5 66.7%), while Artificial Analysis's Coding Agent Index has Sol at a SOTA 80 (+2.8 over Fable 5) and Terminal-Bench 2.1 give...
Cohere launched North Mini Code on June 9 under Apache 2.0, its first developer-focused model. The shape is the pitch: 30B parameters, mixture-of-experts, only ~3B active, and it runs on a single H100. It scores 33.4 on the Artificial Analysis Coding Index, competes on SWE-Ben...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.