Skills
Poisoning RAG context drops accuracy from 77.9% to 43.5%, and number swaps only bite once they hold the majority
A factorial sweep of 588 runs poisoned zero to three of three retrieved passages for Llama 3.1 8B on a FEVER-derived fact-checking task, using entity swap, number swap and negation. Clean accuracy of 77.9% fell to 43.5% with all three passages corrupted, entity swap flipped the largest share of previously correct answers, and number-based corruption stayed flat while poisoned passages were a minority then jumped once they formed a majority. Notably the model mostly abstained rather than inventing new falsehoods, and a lexical-overlap proxy for unsupported generation fell under attack rather than rising. The authors flag coarse automated labels and treat the strategy contrasts as suggestive.
↳ Follow the thread