Fetching from the wire…
Public story · 2026-08-04 · high
The attack watches what the agent learns, then fills the gaps with harmless-looking tasks that combine into a jailbreak.
Why now: The paper posted to arXiv in August 2026, while its authors argue agent teams need to rethink per-write memory filters.
EvoBreak jailbreaks a self-evolving AI agent without writing a single malicious memory, per a paper posted to arXiv (2608.01759). Self-evolving agents update their own memory from what they learn on the job, and EvoBreak turns that update process against them.
That breaks the core assumption behind agent memory safety tools, that scanning each record alone is enough. A filter that checks writes one at a time never sees a jailbreak built from several benign records.
The attack watches what the victim distills from its own experience, then finds the gaps in target-relevant knowledge it hasn't picked up yet.
EvoBreak closes those gaps through ordinary-looking tasks engineered to leave behind exactly the missing pieces. Nothing about any single task reads as malicious. Only the final query, built to activate the accumulated pieces together, does the damage, per the paper.
The attacker never needs direct access to the agent's memory, per the paper. Ordinary-looking tasks are enough to shape what the agent stores on its own.
Each link below shares sources, entities, or timing with this story.
Self-hosted agents read and write their own memory and config to function, which means an attacker can compromise one entirely through legitimate OS system calls with no exploit involved (arXiv 2607.17986). The paper builds a 23-cell attack matrix across Target, Mechanism, Gra...
Poisoned entries in persistent memory force unintended tool selection during retrieval — even against explicit user instructions. Unlike prompt injection targeting input, MCFA targets the memory store, making it persistent and harder to detect. If your agent has long-term memo...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
arXiv 2608.26733 presents an execution-only attack that reconstructs a hosted agent skill without ever asking the victim to reveal it, submitting crafted but ordinary tasks whose results discriminate between candidate hidden behaviors. At the weakest access level, final respon...
The attack needs no instruction, trigger, or retriever optimization, just plainly worded false assertions generated in one pass against a LongMemEval corpus. A four-stage screening pipeline that reaches 0.832 recall on indirect prompt injection rejected none of the poisoned me...
arXiv 2608.05144 runs Manager, Planner, Engineer and Reviewer roles over persistent project state with *fixed* model weights, self-evolving through runtime state and control policies rather than training. 76.8% on AARRI-Bench, and mature waves use 21% fewer solve-input tokens...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.