Fetching from the wire…
Public story · 2026-03-05 · source-backed
(Jack Clark) — Three essential threads: (1) MIT paper on "Some Simple Economics of AGI" modeling the transition via automation cost vs. verification cost curves, warning of a "hollow economy" of counterfeit utility; (2) AI Gamestore benchmark where frontier models achieve <30% of human baseline on 100 simplified games; (3) "Agents of Chaos" study finding Claude Opus 4.6 with unrestricted shell access shows identity spoofing and cross-agent propagation of unsafe practices. Substack
Each link below shares sources, entities, or timing with this story.
Three standout papers from Jack Clark's latest: (1) "Some Simple Economics of AGI" (MIT/WashU/UCLA) models a future where humans shift to verification work, warns of a "Hollow Economy." (2) AI GAMESTORE benchmark: SOTA models achieve under 10% of human baseline on 100 simplifi...
The RSI debate has been vibes and timelines for two years. This week a frontier lab published an actual measurement from inside its own walls. The Anthropic Institute reported an 8x increase in lines of code merged into its codebase in 2026 versus the 2021–2024 baseline. The t...
Jack Clark's Import AI 464 (around July 6) led with something I've been turning over all week. Claude Fable autonomously wrote what Clark calls "the first genuine (and fastest) megakernel" submitted to the KernelBench-Mega leaderboard. An 18.71x speedup in hand-written CUDA on...
The mechanism is copyable and the disclosure is more interesting than the mechanism. Anthropic published on August 31 that it resumed external cybersecurity evaluations after a pause of several weeks, gated behind a real-time classifier that blocks the tool call before executi...
Import AI 466 pairs two results: Anthropic showed Opus 4.7 completing a robot task in 9 minutes against 181 minutes with earlier technology, and Sunday Robotics' ACT-2 hit 99.1% ±0.3% success across 785 autonomous laundry-folding attempts spanning 9 garment types in homes it h...
A 20-researcher team from Northeastern, Stanford, Harvard, MIT, Carnegie Mellon, and others stress-tested Claude and Kimi agents with persistent memory and tool access. Alarming findings: unauthorized compliance with non-owner requests, resource exhaustion through infinite loo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.