Fetching from the wire…
Research2026-07-09 · source-backed
NVIDIA open-sourced its ASPIRE robot skill library on July 1, and Fan framed it as a paradigm shift: robots accumulate reusable skills through repeated failure-and-repair instead of gradient descent, and a dual-arm handover task jumped from 20% to 92% success as the library grew. The framing travels straight into coding agents. It's the same bet that a persistent, expanding skill library, not just bigger weights, is the next unit of capability. This is the clearest primary-source articulation of the idea I've seen, and it's the thread connecting ASPIRE, Beads, and every "harness that learns over time" project in this issue.
Each link below shares sources, entities, or timing with this story.
NVIDIA's GEAR Lab, with CMU and UC Berkeley, released a closed-loop system where coding agents reset physical scenes, run hardware trials, verify outcomes, and rewrite code until a policy works. Jim Fan calls it "AutoResearch in the physical world." Agent teams hit 99% pass@8...
This one annoyed me, in the good way. Researchers took 206 real developer-agent sessions from 13 developers, extracted each developer's preferences from their actual interaction traces via rule-based bootstrapping plus evidence-grounded refinement, then replayed everything aga...
paddo.dev makes the most contrarian read: the letter's substance isn't openness but paragraph nine, defending distillation as "a widely used technique for model improvement" and urging policymakers against "conflating legitimate model development techniques with misappropriati...
The comparison is against GB300 NVL72, with 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark using real recorded agentic coding sessions with context growth, tool calls and sub-agent spawning preserved (NVIDIA). DeepSeek V4 Pro and Qwen3.5 were...
Three frontier models shipped in a single week this month, and teams with a standing eval harness had a routing decision in hours. Anthropic's own agent-eval guidance says 20-50 tasks drawn from your real usage and real failures is enough to detect issues (DeepEval). DeepEval...
xAI launched it July 8, describing it as Opus-class but faster and more token-efficient, at $2/1M in and $6/1M out. Trained across tens of thousands of NVIDIA GB300 GPUs with RL over hundreds of thousands of multi-step software engineering tasks, and trained *alongside Cursor*...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.