Fetching from the wire…
Public story · 2026-02-24 · source-backed
The ARC-AGI-2 breakthroughs reveal a concrete architectural pattern. Symbolica's Agentica achieves 85.28% (with Opus 4.6) using recursive delegation where sub-agents spawn sub-agents, each receiving only relevant state and avoiding context rot. Average 2.6 agents per task, max 9 recursive depth, at $6.94/task. Poetiq's iterative refinement is LLM-agnostic, working across OpenAI/Anthropic/xAI. For Claude Code users: when using agent teams, give each sub-agent only the context it needs. Scoped context beats long windows.
Each link below shares sources, entities, or timing with this story.
Poetiq (6-person ex-DeepMind startup) hit 54% on ARC-AGI-2 at $30.57/task, beating Google's Gemini 3 Deep Think (45%, $77.16/task) through iterative refinement loops — no fine-tuning required. Meanwhile, Symbolica's Agentica reached 85.28% using recursive sub-agent delegation...
poetiq.ai — Detailed technical breakdown of how a $40K-hardware startup achieved 54% on ARC-AGI-2 (vs. Google's 45% at nearly 3x the cost). The key innovation: "learned test-time reasoning" — an iterative refinement meta-system where solutions are generated, receive structured...
The ARC Prize Foundation dropped ARC-AGI-3 on March 25 and the results broke my mental model of how AI capability scales. Symbolica's Arcgentica framework scored 36.08% (113 of 182 playable levels, 7 of 25 games completed) using Claude Opus 4.6 as its backbone. Cost: $1,005. F...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.