Fetching from the wire…
Public story · 2026-03-03 · source-backed
ARC-AGI-3 makes a fundamental shift: instead of static puzzle-solving, it measures agency — a model's capacity to set and pursue goals independently in interactive environments. Public release March 25. If frontier models still fail at ARC-AGI-3 despite succeeding at coding tasks, it validates Chollet's thesis that current LLMs lack genuine generalization. This will become the standard for measuring whether "agentic" AI is real or orchestrated pattern-matching. ARC Prize
Each link below shares sources, entities, or timing with this story.
Chollet confirmed ARC-AGI-3 launches March 25 — the first major format change since 2019. The key shift: ARC-3 tests interactive reasoning and agency — a model's capacity to set and pursue goals independently. The scoring metric provides the first formal human vs. AI action ef...
First major format change since ARC was introduced in 2019. Version 3 tests interactive reasoning and agency — an AI's capacity to set and pursue goals independently. New metric formally compares human vs AI "action efficiency." Developer Toolkit released, public launch March...
The ARC Prize Foundation is launching ARC-AGI-3, the first major format change since 2019. Unlike versions 1 and 2 (static visual reasoning), version 3 uses game-like environments where agents explore without instructions, discover rules, and adapt to hand-crafted levels that...
Drops static puzzles for interactive games where agents must figure out rules through exploration. First format change since 2019. Preview: most humans beat the games easily; AI agents struggle. ARC Prize
- Source: arcprize.org - Date: 2026-02-10 First interactive reasoning benchmark: agents navigate video-game-like environments with no instructions, discovering rules across 1,000+ levels in 150+ hand-crafted environments. Directly challenges "scale is all you need" by requirin...
First interactive reasoning benchmark. Top agent (StochasticGoose) scored 12.58% vs. humans. "Intelligence is efficiency." Agents struggle to convert environmental feedback into coherent strategies. Full launch March 25. ARC Prize
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.