Fetching from the wire…
Top 5 · 2026-02-24 · source-backed
Poetiq (6-person ex-DeepMind startup) hit 54% on ARC-AGI-2 at $30.57/task, beating Google's Gemini 3 Deep Think (45%, $77.16/task) through iterative refinement loops — no fine-tuning required. Meanwhile, Symbolica's Agentica reached 85.28% using recursive sub-agent delegation with Opus 4.6, averaging 2.6 agents per task with max 9 recursive depth. The builder takeaway: invest in your agent orchestration architecture, not just chasing the latest model. Scoped sub-agent context beats cramming everything into one long window.
Each link below shares sources, entities, or timing with this story.
The ARC-AGI-2 breakthroughs reveal a concrete architectural pattern. Symbolica's Agentica achieves 85.28% (with Opus 4.6) using recursive delegation where sub-agents spawn sub-agents, each receiving only relevant state and avoiding context rot. Average 2.6 agents per task, max...
poetiq.ai — Detailed technical breakdown of how a $40K-hardware startup achieved 54% on ARC-AGI-2 (vs. Google's 45% at nearly 3x the cost). The key innovation: "learned test-time reasoning" — an iterative refinement meta-system where solutions are generated, receive structured...
The ARC Prize Foundation dropped ARC-AGI-3 on March 25 and the results broke my mental model of how AI capability scales. Symbolica's Arcgentica framework scored 36.08% (113 of 182 playable levels, 7 of 25 games completed) using Claude Opus 4.6 as its backbone. Cost: $1,005. F...
OpenAI shipped native computer use in GPT-5.4, scoring 75.0% on OSWorld-Verified vs. the 72.4% human baseline (up from 47.3% in GPT-5.2). This is the first general-purpose model to surpass human performance on real desktop workflows. With 1M token context, 92.8% GPQA Diamond,...
Google released Gemini 3.1 Pro on February 19, the first ".1" increment in Gemini's history. The standout metric: 77.1% on ARC-AGI-2, more than double the reasoning performance of Gemini 3 Pro. VentureBeat calls it "Deep Think Mini" — adjustable reasoning depth on demand. Feat...
GPT-5.4 scores 0.26%. Opus 4.6 scores 0.25%. Grok-4.20 scores 0.00%. Humans score 100%. The Decoder covered the ARC-AGI-3 launch on March 25, and the results make every "AGI is here" claim look premature. François Chollet launched ARC-AGI-3 at Y Combinator HQ alongside a fires...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.