Fetching from the wire…
Agents2026-08-15 · source-backed
Holding retrieval, target state, model, decoding and tool budget fixed, researchers compared how a retrieved memory gets used. A target-bound note recording a reusable procedure, bindings to recover, applicability conditions and verification requirements hit 62.3% average success across WebArena, WorkArena and AppWorld. Injecting the full trajectory did worse and degraded sharply as traces grew. arXiv 2608.12847 Store the lesson, not the transcript.
Each link below shares sources, entities, or timing with this story.
Store reusable procedures plus bindings, applicability conditions, and verification requirements instead of the raw trace. That gained 10.7 points of success on WebArena, WorkArena, and AppWorld while cutting online tokens 48.9% (arXiv). Direct trace reuse gets worse as your t...
The pattern most agent memory systems use, dumping a retrieved trajectory into context, degrades as traces lengthen and as source-specific values diverge from the target. QCR replaces the dump with a structured memory holding reusable procedures plus bindings, applicability co...
SkillForge (arXiv 2608.24747) notes that skill-extraction approaches like SkillRL never verify whether a stored skill still works against the current environment, so the bank grows monotonically while quality rots. It makes skill usage explicit during interaction so RL optimiz...
Numbers first, because they're the whole argument. On AppWorld's 168 tasks with DeepSeek-V3.2: ALTK-Evolve hit 89.3% goal completion at 263K tokens per task. ACE hit 80.4% at 634K (IBM Research on HF). Nine points better, 41% of the tokens. On the weaker gpt-oss-120b: 56.0% at...
arXiv 2608.04755 injected Android permission popups into real GUI tasks across four frontier multimodal LLMs with synchronized screenshots and UI trees. Holding the task fixed and changing only the requesting app flipped grants from 26/32 to 0/32, an App-Trust Bias. Holding th...
arXiv 2607.12227 (Wang et al., incl. Hajishirzi, Tsvetkov, Dasigi) finds two methodological holes in the self-improving-agent literature: methods are never compared against simpler baselines at matched compute budgets, and final performance gets reported on the same public ben...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.