Fetching from the wire…
Agents2026-08-11 · source-backed
arXiv 2608.08311 describes an agent that continuously rewrites its own tools, prompts, context assembly, and core implementation through reviewed commits. On Opus 5: 86.74% Terminal-Bench 2.1, 90.69% OSWorld-Verified, 0.2301 normalized reward on CL-Bench. It also documents "Hope," a 161-day public deployment evolving freely across seven surfaces. Roman Yampolskiy, best known for arguing advanced AI is uncontrollable, is a co-author, which is either reassuring or the opposite. 1,060 upvotes on HuggingFace Daily Papers, about 13x the next paper that day.
Each link below shares sources, entities, or timing with this story.
A team at UC Berkeley RDI built an automated scanning agent that achieved near-perfect scores on eight major AI agent benchmarks. SWE-bench Verified: 100%. Terminal-Bench: 100%. WebArena: approximately 100%. FieldWorkArena: 100%. GAIA: roughly 98%. OSWorld: 73%. The agent didn...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Launch HN from YC S26 founders (ex-AppLovin and Citadel, after six pivots) pitches a speed-focused harness rather than a model: model routing, targeted code search instead of whole-repo embedding, context management, and turn batching they say cut round trips 16% and costs 27%...
Same weights. Same tasks. Same time budget. One change to the harness, and fail-to-pass on SWE-bench Verified went from 28% to 49%. The setup in arXiv 2608.26218 is clean enough to be worth describing precisely. 169 SWE-bench Verified tasks, a 20,480-token context window, a fi...
June leaderboards increasingly score a weighted blend of Terminal-Bench 2.0, BrowseComp, and OSWorld-Verified instead of SWE-bench alone, with BenchLM weighting "agentic" browse-and-do workflows at 22%, its highest single category. When you evaluate a tool, the headline chat o...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.