Fetching from the wire…
Agents2026-07-13 · source-backed
This method retains four categories of reusable context (task specs, data schemas, tool configs, output constraints) while discarding session-specific reasoning, enabling role-based workspace transfer across users (arXiv:2607.09493). It reports 96% completion versus 79% without memory, 14x faster zero-token refresh for recurring updates, and 97x lower per-invocation token cost versus injecting raw data each time. This is the concrete blueprint for what "agent memory" should actually persist. Not chat history. The reusable scaffolding.
Each link below shares sources, entities, or timing with this story.
This is the most directly usable paper of the day and it does something rare: it bolts onto an existing agent without touching it. Ledger is a deterministic runtime wrapper that distills an agent's completed interactions into explicit state. What has been observed, what has be...
book-to-skill (14,194 stars, +595 today) compiles a PDF into a ~4K-token SKILL.md plus ~1K-token per-chapter files loaded on demand, reporting 24-51x fewer tokens than dumping the book and ~5K tokens resident versus 119K-256K for a context dump, at roughly $1 per book to conve...
263 task pairs across 42 executable sandbox environments, where each pair holds one task the agent should complete and a near-identical one it should refuse or escalate. Every agent benchmark I've seen measures completion rate. This measures the inverse, and the inverse is the...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
Recuris (arXiv 2608.24876) keeps a Working Memory tracking current task progress separate from an Experiential Memory of learned skills, so skill selection indexes against what the task needs now rather than the whole history. It improves 35 of 37 model-benchmark pairs, gains...
arXiv 2608.05906 keeps a dual-polarity memory of verified corrections and observed dead ends for Text-to-SQL repair: 66.34% to 69.79% on Spider, 47.35% to 48.44% on BIRD. Then the authors say the quiet part: paired analysis supports the Spider gain but is weak on BIRD, MERIT i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.