Fetching from the wire…
Public story · 2026-07-27 · high
Anthropic cut a robot task from 181 minutes down to 9 with Opus 4.7. Jack Clark says robotics is now accelerating on someone else's scaling curve.
Why now: Clark draws the comparison in Import AI 466, part of the July 27 briefing coverage.
Sunday Robotics' ACT-2 folded laundry correctly 99.1% of the time across 785 attempts in homes it had never seen, per Import AI 466. It handled nine garment types with zero home- or garment-specific tuning.
The 99.1% figure matters because manual tuning is what's kept household robots stuck in demos, never products. Zero-shot transfer across new homes and garments is the difference between the two.
Anthropic's own numbers point the same way. Opus 4.7 completed a robot task in 9 minutes, down from 181 minutes on earlier technology, according to the same newsletter.
Sunday's recipe, per Import AI: scale pretraining, then hill-climb with a small amount of in-house data. Sunday found that as its pretrained model gets stronger, gains from that small dataset get more transferable, not less.
Clark's read: the bitter lesson has arrived for robotics. Sunday rode a general-purpose model up its scaling curve, doing little robot-specific engineering on top. If he's right, the next leap in home robots comes from the next foundation model release, not new robotics research. Worth watching whether Sunday's transfer gains keep climbing as pretrained models get stronger, or hit a wall that more compute doesn't fix.
Each link below shares sources, entities, or timing with this story.
The RSI debate has been vibes and timelines for two years. This week a frontier lab published an actual measurement from inside its own walls. The Anthropic Institute reported an 8x increase in lines of code merged into its codebase in 2026 versus the 2021–2024 baseline. The t...
Jack Clark's Import AI 464 (around July 6) led with something I've been turning over all week. Claude Fable autonomously wrote what Clark calls "the first genuine (and fastest) megakernel" submitted to the KernelBench-Mega leaderboard. An 18.71x speedup in hand-written CUDA on...
A group including Tri Dao, Jürgen Schmidhuber, Joshua Tenenbaum, Thomas Griffiths, and James Whittington encoded 70 hidden-rule discovery games as short strings where both the transformation rules and the win conditions must be inferred through experimentation (arXiv 2608.1259...
In Import AI 463, Jack Clark leads with ENPIRE, software that puts real-world robotics into autonomous experiment-and-execute cycles, the embodied version of how software agents self-improve. The agentic self-improvement pattern is being pushed into physical systems. That's th...
The core team released lemans on August 24 after deciding the Ruby community shouldn't have to run Python-based Harbor, and benchmarked four models on 63 Rails tasks (Rails). ox-alpha 52/63; Terra 49/63 at $0.20 and a 182-second median; open-weight Qwen 3.8-27B 48/63 but at a...
Import AI 470 walks through a METR note measuring where LLMs moved the needle (Import AI). Vulnerability reports accelerated across cURL, OpenSSL, Firefox, Microsoft, the US NVD and OSV in 2026 against 2025. Math shows arXiv submissions doubling in some areas plus named result...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.