Fetching from the wire…
Agents2026-08-13 · source-backed
arXiv 2608.11632 argues storage retention doesn't identify authoritative state: unmediated updates by models, tools, and background workers cause stale overwrites and self-authorizing privilege escalation. Untrusted components propose typed changes against an exact predecessor head; a short activation transaction revalidates ownership, freshness, and effect uniqueness before recording exactly one disposition. Verified across 2,808,230 reachable states and 5,526,474 transitions with zero invariant violations. Formal methods showing up in agent memory design is new and I want more of it.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.13292 characterized 28 APR approaches on SWE-bench Verified: median 121.78% more total changes, 80.91% more net changes, 43.99% higher cyclomatic complexity than the developer patch, even when correct. The verbosity is rooted in capability-oriented design and resist...
The study extracted 130 clean atomic state transitions from 707 real issues in SWE-bench Lite and Verified. Plain RAG scored 0.57-0.59 answer accuracy; an LLM reranker didn't help and added latency, about 18 seconds against 2.1. A (subject, relation, object) supersession memor...
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
RETRACE has a verifier infer what problem the patch appears to solve using only the patch and trajectory, then compares that inference against the real issue. Training-free, lifted Pass@1 by 7.0% and 3.6% on mini-SWE-agent over SWE-bench Verified. The information-hiding trick...
arXiv 2608.03609 formalizes agentic systems over relational data as Stateful Tool-Enabled Agentic Deployments and proves verification against First-Order CTL specs is undecidable. Under a finite-domain restriction it becomes PSPACE-complete, but only if renaming opaque identif...
Equips code agents with structured memory built from historical commits, distilling intent-to-code mappings with self-refinement via verification feedback. Using DeepSeek-V3.2 as backbone, boosts SWE-bench Verified from 68.4% to 77.8% — new SOTA. Co-evolution with project hist...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.