Fetching from the wire…
Public story · 2026-02-23 · source-backed
Each link below shares sources, entities, or timing with this story.
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top thre...
36. arXiv — CUWM 37. arXiv — IntentCUA 38. arXiv — SpargeAttention2 39. arXiv — Arcee Trinity Large 40. ARC Prize — ARC-AGI-2 41. Arcee Blog ---
I've been saying for months that the real gains aren't in switching models. They're in how you set up the environment around the model. Now there's quantitative proof. Stanford IRIS Lab published Meta-Harness, a system that autonomously evolves its own coding harness, system p...
Anthropic shipped cross-session messaging for Claude Code on August 7, macOS and Linux, version 2.1.224 or higher. Two new tools: ListAgents discovers other active sessions on your machine, SendMessage delivers text to one by name. Messages between sessions on the same machine...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
Chris Lattner's CCC Review — The definitive expert assessment of AI coding capabilities. Required reading. Source: Modular Blog Boris Cherny on Lenny's Podcast — Internal Anthropic metrics, the "constraints + unlimited tokens" formula, and why "software engineer" as a title go...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.