Fetching from the wire…
Models2026-08-27 · source-backed
Quesma ran the model across GPQA Diamond, IFBench and Terminal-Bench 2.1 (89 agentic coding tasks) on L40S, H100 and H200 via Modal. Q4_K_M at 17 GB matched BF16 at 55 GB within a point on all three. UD-Q2_K_XL at 10.7 GB held instruction-following but dropped Terminal-Bench from about 77% to about 72%. Both 1-bit quants at 6.2 GB fell to 15-20% on GPQA Diamond, which is around random chance, and degraded further on longer reasoning (Quesma). Q4_K_M remains the default answer and this is the cleanest evidence for it I've seen this month.
Each link below shares sources, entities, or timing with this story.
The 397B MoE scores 86.1 on Terminal-Bench 2.1 (Terminus-2) against Claude Opus 4.8's 85.0, 86.0 on SWE-bench Verified, and 92.8 on GPQA Diamond. Hugging Face But it trails badly on the harder agentic rows: 13.5 versus 21.1 on Frontier-Bench v0.1, 59.5 versus 69.7 on NL2Repo....
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Cohere launched North Mini Code on June 9 under Apache 2.0, its first developer-focused model. The shape is the pitch: 30B parameters, mixture-of-experts, only ~3B active, and it runs on a single H100. It scores 33.4 on the Artificial Analysis Coding Index, competes on SWE-Ben...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
Alibaba's most powerful model dropped the same day as K2.6. First on SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, and three others. 260K context window, API compatible with both OpenAI and Anthropic specs.
The 397B MoE scores 86.1 on Terminal-Bench 2.1, 86 on SWE-Bench Verified, 56 on DeepSWE, 92.8 on GPQA Diamond and 44.6 on HLE, which the team frames as comparable to Claude Opus 4.8. The method is a closed self-improvement loop where the model proposes its own tasks and scaffo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.