Fetching from the wire…
Public story · 2026-02-20 · source-backed
Google released Gemini 3.1 Pro on February 19, the first ".1" increment in Gemini's history. The standout metric: 77.1% on ARC-AGI-2, more than double the reasoning performance of Gemini 3 Pro. VentureBeat calls it "Deep Think Mini" — adjustable reasoning depth on demand. Features a 1M token context window. Critical business detail: same pricing as Gemini 3 Pro, effectively a free performance upgrade for API users. Available across Gemini app, NotebookLM, API, Vertex AI, and Gemini CLI.
The frontier model landscape now has three competitors with production-grade reasoning: Claude Opus 4.6 (80.8% SWE-bench), Gemini 3.1 Pro (77.1% ARC-AGI-2), and GPT-5.3 (paused). The competitive axis is shifting from raw benchmark scores to practical developer integration and adjustable reasoning depth.
Each link below shares sources, entities, or timing with this story.
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
The biggest single-benchmark jump in a frontier model update: 37.6% → 68.8% on ARC-AGI-2. For comparison, GPT-5.2 scored 54.2%, Gemini 3 Pro 45.1%. ARC-AGI-2 measures novel problem-solving on adversarially constructed tasks. A near-doubling suggests genuine reasoning improveme...
The ARC Prize Foundation dropped ARC-AGI-3 on March 25 and the results broke my mental model of how AI capability scales. Symbolica's Arcgentica framework scored 36.08% (113 of 182 playable levels, 7 of 25 games completed) using Claude Opus 4.6 as its backbone. Cost: $1,005. F...
The LogRocket February 2026 Power Rankings reshuffled: | Rank | Tool | Price | Key Feature | |------|------|-------|-------------| | 1 | Windsurf | $15/mo | Arena Mode + parallel worktrees | | 2 | Antigravity (Google) | Free-$250 | Deep Google ecosystem integration | | 3 | Cur...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top thre...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.