Fetching from the wire…
Agents2026-08-27 · source-backed
A single-author study models supersession in inherited agent memory: a constraint true when written, since withdrawn by a newer authoritative record. Given a two-record verification budget, agents checked the provenance path in about one episode in five and produced stale-consistent decisions in 77.3%, 74.7% and 74.7% of episodes across a primary run, a fresh-wording replication and a held-out domain (arXiv 2608.25553). Re-assigning one of the two slots to the critical provenance path raised current-record-consistent decisions by 74.0, 72.7 and 61.3 points, positive in six of six models, and changed nothing when the record agreed with memory. Same budget, different allocation.
Each link below shares sources, entities, or timing with this story.
Same weights. Same tasks. Same time budget. One change to the harness, and fail-to-pass on SWE-bench Verified went from 28% to 49%. The setup in arXiv 2608.26218 is clean enough to be worth describing precisely. 169 SWE-bench Verified tasks, a 20,480-token context window, a fi...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
SpecPath (arXiv 2608.09799) holds repository, final contract, verifier, agent system, and execution budget constant, varying only the revision history by which a spec became final. Same end state, different path there. Thirty-five percent failed on at least one contract-equiva...
Seven models. Five harnesses. Controlled fact-withholding with injected faults. arXiv 2608.16630 is the most operationally direct paper I've read on harness design, and it produces three results that each change what I do this week. One: availability decides outcomes, not dist...
Numbers first, because they're the whole argument. On AppWorld's 168 tasks with DeepSeek-V3.2: ALTK-Evolve hit 89.3% goal completion at 263K tokens per task. ACE hit 80.4% at 634K (IBM Research on HF). Nine points better, 41% of the tokens. On the weaker gpt-oss-120b: 56.0% at...
Anthropic shipped cross-session messaging for Claude Code on August 7, macOS and Linux, version 2.1.224 or higher. Two new tools: ListAgents discovers other active sessions on your machine, SendMessage delivers text to one by name. Messages between sessions on the same machine...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.