Fetching from the wire…
Models2026-08-20 · source-backed
The 397B MoE scores 86.1 on Terminal-Bench 2.1, 86 on SWE-Bench Verified, 56 on DeepSWE, 92.8 on GPQA Diamond and 44.6 on HLE, which the team frames as comparable to Claude Opus 4.8. The method is a closed self-improvement loop where the model proposes its own tasks and scaffolds for RL rollouts, jointly optimized with GRPO instead of training on human-curated task sets. (Ornith) But a top r/LocalLLaMA commenter's side-by-side shows the 35B-A3B losing to Qwen3.8-27B on Terminal-Bench (68.5 vs 73.0) and DeepSWE (22.0 vs 42.2) while winning NL2Repo (46.2 vs 42.3). So the headline doesn't transfer to the sizes most people can actually run. Worth noting the tension with arXiv 2608.19072 above: self-generated training curricula at the model level, while agents doing model-level strategy selection can't revise their own approach.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The 397B MoE scores 86.1 on Terminal-Bench 2.1 (Terminus-2) against Claude Opus 4.8's 85.0, 86.0 on SWE-bench Verified, and 92.8 on GPQA Diamond. Hugging Face But it trails badly on the harder agentic rows: 13.5 versus 21.1 on Frontier-Bench v0.1, 59.5 versus 69.7 on NL2Repo....
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.