Research
ToxicRAG Poisons RAG With One 'Knowledge Update' Document per Target at 61-91% Attack Success
ToxicRAG (arXiv 2609.11082) writes each poisoned document as a narrative: it acknowledges the accepted answer, invents events that overturn it, and cites fake authorities for the attacker's answer. With one document per target question, it hit 0.61-0.91 attack success across Natural Questions, HotpotQA and MS-MARCO with four victim LLMs and four dense retrievers, matching or beating the strongest baseline in all 12 combinations. Filters that look for documents asserting a bare answer will miss this style.
Source
↳ Follow the thread