Fetching from the wire…
Research2026-08-18 · source-backed
A group including Tri Dao, Jürgen Schmidhuber, Joshua Tenenbaum, Thomas Griffiths, and James Whittington encoded 70 hidden-rule discovery games as short strings where both the transformation rules and the win conditions must be inferred through experimentation (arXiv 2608.12593). All 70 solved by at least one human on first attempt. 21 public, 49 held private. Per Jack Clark's Import AI 469, Opus 5 and Fable 5 led, and only those two cleared any Tier-7 task at 0.2 success. Clark predicts human parity by mid-2027 and treats that as the gate on recursive self-improvement.
Each link below shares sources, entities, or timing with this story.
Jack Clark's Import AI 464 (around July 6) led with something I've been turning over all week. Claude Fable autonomously wrote what Clark calls "the first genuine (and fastest) megakernel" submitted to the KernelBench-Mega leaderboard. An 18.71x speedup in hand-written CUDA on...
Import AI 466 pairs two results: Anthropic showed Opus 4.7 completing a robot task in 9 minutes against 181 minutes with earlier technology, and Sunday Robotics' ACT-2 hit 99.1% ±0.3% success across 785 autonomous laundry-folding attempts spanning 9 garment types in homes it h...
The RSI debate has been vibes and timelines for two years. This week a frontier lab published an actual measurement from inside its own walls. The Anthropic Institute reported an 8x increase in lines of code merged into its codebase in 2026 versus the 2021–2024 baseline. The t...
The mechanism is copyable and the disclosure is more interesting than the mechanism. Anthropic published on August 31 that it resumed external cybersecurity evaluations after a pause of several weeks, gated behind a real-time classifier that blocks the tool call before executi...
The extension takes autonomous actions (reading, clicking, form-filling) without per-action approval, gated by a safety classifier validating each action plus probes scanning page content for injection attempts. Anthropic publishes per-model attack success rates with safeguard...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.