Giving the Memory Curator Read-Only World Tools Takes CLBench From 39% to 73% and Halves Task Cost
Post-task memory curators that only see completed trajectories preserve errors, overgeneralize partial evidence, and retain stale facts. Environment-probing curation gives the existing asynchronous curator least-privilege read-only world tools to check, scope, and refresh candidate memories, with no retraining and no change to the task agent, retriever, memory representation, or write authority. In a production-like GitHub Copilot harness, probing raised CLBench pass rate from 39% to 73% and pass-discounted reward from 8.60 to 22.60 while cutting queries from 8.8 to 4.7 per question and task-agent cost from $3.38 to $1.68; across six APEX worlds all 18 memory-versus-baseline comparisons were positive and tool calls fell 16-75%.
↳ Follow the thread