Fetching from the wire…
Vibe Coding2026-08-06 · source-backed
Open-sourced under MIT, built on two abstractions: a Recursive Language Model (a persistent IPython REPL where the model calls sub-agents as functions, keeping programmatic access to context and history) and a Continual Harness treating its own prompts, skills, memory and sub-agents as CRUD-able state, with a /refine command that analyzes trajectories and makes targeted harness edits. Also working Game Boy Color and SEGA Genesis emulators on EmulatorBench. The ARC number is self-reported and explicitly unendorsed by ARC Prize. Ignore the number, copy the harness-as-editable-state design.
Each link below shares sources, entities, or timing with this story.
If you're tuning an agent system, upgrade the model doing the tuning before you rewrite a single line of the harness. That's the finding from HarnessOpt-Bench (arXiv 2608.06301, Scale AI), which tests whether frontier models can improve an agent *system* rather than write code...
GPT-5.4 scores 0.26%. Opus 4.6 scores 0.25%. Grok-4.20 scores 0.00%. Humans score 100%. The Decoder covered the ARC-AGI-3 launch on March 25, and the results make every "AGI is here" claim look premature. François Chollet launched ARC-AGI-3 at Y Combinator HQ alongside a fires...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Everyone is building summarize-and-evict context management. Compaction, rolling summaries, hierarchical memory, vector-store recall. The entire agent-memory category assumes the answer is to throw away history intelligently. PRO-LONG (arXiv 2607.20064) keeps the complete stru...
Two weeks ago the US government forced Anthropic to pull Mythos 5 offline under an emergency export directive, on the theory that frontier cyber capability is dangerous enough to gate. This week a Chinese lab released a model you can download under an MIT license that benchmar...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.