Agents
Latent Adversarial Detection: Activation Probing Catches Multi-Turn Prompt Injection at 93.8% Accuracy
Researchers discovered 'adversarial restlessness' — a characteristic activation path length in LLM residual streams that emerges as multi-turn attacks progress through trust-building, pivoting, and escalation phases. Using five scalar trajectory features, detection accuracy jumps from 76.2% to 93.8% on synthetic data and reaches 89.4% with 2.4% false positive rate on combined training. The signal replicates across four model families (24B–70B parameters). Three-phase turn-level labels proved essential; binary labels produced 50–59% false positives, explaining why prior work underperformed.
Source
↳ Follow the thread