Fetching from the wire…
Agents2026-08-27 · source-backed
The method synthesizes security skills offline from known attacks and recorded agent failures, injects them into the system prompt at session start, and leaves them active through the tool-use loop. Across six models on RedCode the default all-classes skill dropped malware-generation severity from 3.37 to 0.58 and reached a 43.6% execution attack success rate, comparable to Llama Guard 3's 42% (arXiv 2608.25817). No auxiliary classifier, no execution monitor in the trajectory, which makes it available to API-only deployers who can't touch weights. The paper compares three fixed system-prompt budgets, none of which needs runtime request routing.
Each link below shares sources, entities, or timing with this story.
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings and seven binary classifiers trained on 30,568 samples, reporting 95.9% recall end-to-end against 255ms for Llama Guard 2 and 723ms for SafeDecoding, with 100% detection of GCG suffix...
arXiv 2607.05120 defines a new attack class: instead of hijacking what the agent does, corrupt which resources it acts upon, by injecting probabilistic delimiters that make untrusted data read as trusted metadata (resource IDs, data origin, tool-call history). On agents where...
New paper argues both regex and fine-tuned classifiers share critical failure modes for detecting prompt injection. Regex misses paraphrased attacks, classifiers get bypassed at >50% success rates by adaptive adversaries. The paper proposes seven detection techniques from fiel...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
arXiv 2608.26733 presents an execution-only attack that reconstructs a hosted agent skill without ever asking the victim to reveal it, submitting crafted but ordinary tasks whose results discriminate between candidate hidden behaviors. At the weakest access level, final respon...
arXiv 2608.05604 names the mismatch precisely: current systems retrieve skills as packages but compress them as prose, which destroys the execution contract. SkillZip does contract-preserving compression over section-level graphs, rewriting recurring valid motifs into reversib...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.