Audio prompt injection completes in 49% of cells where images complete in 1%, across 720 runs on six agent frameworks
MMPIBench (arXiv 2609.09404, 2026-09-08) pushed a fixed attack set through six visual carriers — OCR text, overlays, EXIF metadata, QR codes, fake interfaces, hybrids — across 720 runs covering six frameworks, five models and four attacker objectives, tracing each injection from perception through planning to the tool call. Visual attacks were attempted in 12.8% of runs but completed in only about 1%, with nearly the whole gap closed at the planning step, and the model mattered far more than the framework: one model never attempted an attack and flagged the injection 59.7% of the time while two others attempted in 23.6%. Extending to audio flipped the result — only two of five models ingest audio and only three of six frameworks deliver it, but where the signal arrives the attack completed in 49% of cells and 75% for one model. Vision has been hardened by training; the other perceptual channels have not.
Source
↳ Follow the thread