Fetching from the wire…
Public story · 2026-03-11 · source-backed
The first interactive reasoning benchmark: instead of input/output grids, agents face novel games in an ARC grid world where they must discover rules through trial and error, track state, and learn on the fly. Given METR showed SWE-bench PRs aren't mergeable and EsoLang-Bench proved memorization, ARC-AGI-3 could become the gold standard for measuring actual AI reasoning. ARC Prize
Each link below shares sources, entities, or timing with this story.
The ARC Prize Foundation is launching ARC-AGI-3, the first major format change since 2019. Unlike versions 1 and 2 (static visual reasoning), version 3 uses game-like environments where agents explore without instructions, discover rules, and adapt to hand-crafted levels that...
ARC-AGI-3 makes a fundamental shift: instead of static puzzle-solving, it measures agency — a model's capacity to set and pursue goals independently in interactive environments. Public release March 25. If frontier models still fail at ARC-AGI-3 despite succeeding at coding ta...
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage ri...
- Source: arcprize.org - Date: 2026-02-10 First interactive reasoning benchmark: agents navigate video-game-like environments with no instructions, discovering rules across 1,000+ levels in 150+ hand-crafted environments. Directly challenges "scale is all you need" by requirin...
First interactive reasoning benchmark. Top agent (StochasticGoose) scored 12.58% vs. humans. "Intelligence is efficiency." Agents struggle to convert environmental feedback into coherent strategies. Full launch March 25. ARC Prize
First major format change since ARC was introduced in 2019. Version 3 tests interactive reasoning and agency — an AI's capacity to set and pursue goals independently. New metric formally compares human vs AI "action efficiency." Developer Toolkit released, public launch March...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.