Fetching from the wire…
Public story · 2026-08-07 · high
FinEvo-Bench tested 120 real financial cases; rubric feedback beat reference-answer scoring in every scaffold.
Why now: Covered in the August 7 briefing on self-evolving agent research.
Four self-evolving scaffolds beat non-evolving baselines on 120 financial tasks, and in Claude Code, skills alone beat pairing them with memory, per Deng et al.
The stakes: evolution cut compliance issues by up to 0.44 per task, an error rate that compounds fast in regulated finance work.
The study, FinEvo-Bench, ran the four scaffolds against paired non-evolving controls on a shared Qwen3.7-Max backbone across 20 business scenes in six financial domains.
Letta posted the highest evolved score, 91.65, and the fewest compliance issues, 0.09 per task. Codex saw the largest gain from evolution, adding 19.37 points over its non-evolving control.
Rubrics beat reference answers as feedback in every scaffold tested, not just Claude Code. The paper doesn't explain why, leaving open whether graded criteria beats answer-matching outside finance tasks too.
The lesson for builders: don't assume stacking memory onto a skill library is free upside. In Claude Code here, it cost points instead of adding them. Test skill-only against skill-plus-memory on your own tasks before wiring both in by default.
Each link below shares sources, entities, or timing with this story.
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
I've spent real hours tuning the CLAUDE.md in my own repos. Rewriting architecture notes. Adding conventions. Trimming when it got long. So this one stung. arXiv 2607.27250 ran a two-agent ablation across Claude Code and Codex: 17 real tasks from 3 repositories, 288 gold-test-...
A paper from Xiao Yu, Baolin Peng, and Ruize Xu makes a claim that seems obvious once stated and is genuinely new as a training methodology: modern agents are inseparable from their inference harnesses, so training them in stripped-down RL sandboxes produces a train/serve mism...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
SkillSentry (arXiv 2608.09253) targets the gap where an agent completes a task under skill guidance then fails the same task on a repeat run. It defines a DSL for runtime guidance, initializes it from skill specs plus insights mined from historical successful *and failed* trac...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.