Fetching from the wire…
Security2026-04-21 · source-backed
New paper argues both regex and fine-tuned classifiers share critical failure modes for detecting prompt injection. Regex misses paraphrased attacks, classifiers get bypassed at >50% success rates by adaptive adversaries. The paper proposes seven detection techniques from fields outside NLP. If you're building agent pipelines processing untrusted input, pattern matching alone won't save you.
Each link below shares sources, entities, or timing with this story.
arXiv 2609.09553 shows cipher-based covert-communication jailbreaks no longer need fine-tuning on an encrypted corpus. In-context learning is enough, and alignment is significantly weakened or bypassed once the exchange runs through the learned encoding. Demonstrated against m...
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings and seven binary classifiers trained on 30,568 samples, reporting 95.9% recall end-to-end against 255ms for Llama Guard 2 and 723ms for SafeDecoding, with 100% detection of GCG suffix...
The method synthesizes security skills offline from known attacks and recorded agent failures, injects them into the system prompt at session start, and leaves them active through the tool-use loop. Across six models on RedCode the default all-classes skill dropped malware-gen...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
A paper reframing indirect prompt injection as test-time search builds an agentic attacker doing environment reconnaissance, structured strategy reasoning and adaptive evaluation against victim feedback (arXiv 2609.04495). More attacker compute consistently improves both vulne...
arXiv 2608.27141 proves a separation result: against an attack whose evidence is fragmented across iterations, any monitor whose safety state resets each trajectory has a true-positive rate equal to its false-positive rate, no matter how expressive it is. A monitor retaining c...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.