Fetching from the wire…
Public story · 2026-03-19 · source-backed
MiniMax shipped M2.7 on March 18 and made a claim nobody else has made with receipts: the model participated in its own R&D cycle. Not "we used AI to help train it" marketing. MiniMax says M2.7 autonomously handled 30–50% of the development workflow — reading logs, debugging failures, analyzing metrics, and optimizing performance scaffolds — while humans focused on architecture decisions and safety. The result: a model that scored 56.22% on SWE-Pro (matching GPT-5.3-Codex), 55.6% on VIBE-Pro (near Opus 4.6 parity), and hit a 97% skill adherence rate across 40+ complex skills exceeding 2,000 tokens each. The GDPval-AA ELO of 1495 is the highest among open-source models. CnTechPost Latent Space MiniMax
The pricing is where this gets real. At $0.30 input / $1.20 output per million tokens, M2.7 costs roughly one-third of GLM-5 while matching its reasoning benchmarks. Production incident recovery times dropped to under three minutes in real-world engineering scenarios. The model achieved a 66.6% medal rate across 22 ML competitions — the kind of metric that matters because competition submissions are adversarial by nature.
Available today on MiniMax Agent, their API, Ollama, OpenRouter, and Vercel. If you're building agent pipelines and paying frontier-model prices for Sonnet-class reasoning, M2.7 is the first credible alternative where the benchmarks, the price, and the availability all line up simultaneously. The self-evolution angle is the longer-term story — if models can meaningfully participate in their own improvement loops, the gap between releases compresses.
Each link below shares sources, entities, or timing with this story.
I've spent the last year assuming that if I wanted real agentic coding quality, I paid for a closed model. That assumption took a hit on June 1. MiniMax shipped M3 with a new sparse-attention architecture (they call it MSA) that handles up to 1M tokens at roughly 9x prefill an...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
DeepClaude hit 470 points on Hacker News. It swaps Claude Code's API backend to DeepSeek V4 Pro while preserving the full agent loop: file editing, bash execution, git tooling, the whole workflow. DeepSeek V4 Pro scores 96.4% on LiveCodeBench at a fraction of Anthropic's prici...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.