Research
A RAG Poison Split Across Harmless Passages Clears 80% Attack Success and Evades Single-Document Defenses
InceptionRAG (arXiv 2609.16818, submitted 15 Sep 2026) abandons single-point payload injection and instead fragments the payload into a chain of dormant passages that each look benign under isolated inspection but, when retrieved together, lead the model to self-deduce the target misinformation via multi-hop reasoning. Across three datasets and three LLMs it exceeds 80% attack success rate under adversarial constraints while bypassing defenses built for traditional single-document injection. The authors note the paradox directly, stronger reasoning increases vulnerability, and propose a document-isolation defense called HODOR that decouples the adversarial logical dependencies.
↳ Follow the thread