Fetching from the wire…
Agents2026-08-28 · source-backed
arXiv 2608.27141 proves a separation result: against an attack whose evidence is fragmented across iterations, any monitor whose safety state resets each trajectory has a true-positive rate equal to its false-positive rate, no matter how expressive it is. A monitor retaining cross-iteration state separates them perfectly. The obvious patch, a geometrically decaying risk score, fails because the cooling-off period a patient adversary must wait is a constant independent of the horizon. Their LoopHarness keeps a persistent non-decaying loop-level safety state. If your agent monitoring resets per run, it's measuring nothing against an attacker who's willing to be slow.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.27951 separates the capability a model releases from the evidence it has about downstream use, and shows that when that evidence is copyable (a request, a persona, an interaction history an attacker can imitate) there's an exact worst-case floor on attacker assistan...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
Self-hosted agents read and write their own memory and config to function, which means an attacker can compromise one entirely through legitimate OS system calls with no exploit involved (arXiv 2607.17986). The paper builds a 23-cell attack matrix across Target, Mechanism, Gra...
"Adaptive Adversaries" (arXiv:2607.18063) tests agents against attackers that adapt across turns instead of firing one-shot prompts. Claude Opus 4.6 and GPT-5.4 tied at 5.4% aggregate, but per-scenario variance was extreme, with Opus hitting 60% on one scenario where competito...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
The attack needs no instruction, trigger, or retriever optimization, just plainly worded false assertions generated in one pass against a LongMemEval corpus. A four-stage screening pipeline that reaches 0.832 recall on indirect prompt injection rejected none of the poisoned me...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.