Fetching from the wire…
Security2026-08-30 · source-backed
A paper on arXiv shows attack success rate is an attacker-tunable variable, and a reverse-training framework produces low-ASR backdoors that keep clean-input performance while the backdoor behavior stays intact. State-of-the-art defenses fail consistently under low-ASR conditions across multiple datasets, attack families and architectures. (arXiv 2608.27288) Anyone screening third-party model weights is screening for a signal the adversary decides how loud to make.
Each link below shares sources, entities, or timing with this story.
Splitting ASR into covert success, injections leaving no trace in the final response, and overt success, ones a user can spot, follows from the ReAct format where the final response summarizes the most recent action (arXiv 2608.30362). A trace that hands control back to the us...
Agent-agnostic middleware with two halves: System I handles known attacks through a Tier-0 library of rule-based detection scripts backed by Tier-1 optimized LLM inference, System II watches for abnormal signals and attempts to synthesize a new defense. It matches standard-ben...
13 public sources consolidated into 9,740 skills (7,505 malicious, 2,235 benign) across 11 harmonized attack categories. Learned text detectors score 0.882-0.932 Macro-F1 under random splits but collapse to 0.653-0.665 source-disjoint. arXiv Three off-the-shelf skill scanners...
State-corruption attacks work because attacker-controlled data makes false claims about the environment that slip past injection filters, since the text reads like an ordinary tool result. PIPES screens each response unit two ways: static field contracts where a schema gives s...
arXiv 2608.27141 proves a separation result: against an attack whose evidence is fragmented across iterations, any monitor whose safety state resets each trajectory has a true-positive rate equal to its false-positive rate, no matter how expressive it is. A monitor retaining c...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.