Fetching from the wire…
Public story · 2026-03-16 · source-backed
Researchers demonstrate LLMs assign authority based on formatting rather than source, enabling 61% success on agent exfiltration tasks. Novel "role probes" predict attack success before generation begins. No defenses proposed — the gap is fundamental. arXiv
Each link below shares sources, entities, or timing with this story.
MIT Technology Review reported July 30 on "Prompt Injection as Role Confusion," which shows models identify who's speaking from an insecure feature, style, and that "role tags were a formatting trick that became the security architecture" of modern LLMs. If that holds, every h...
Simon Willison's June 22 write-up surfaces "Prompt Injection as Role Confusion" (Ye, Cui, Hadfield-Menell, ICML 2026), which argues injection works because models infer the speaker from a text's *style*, not its labeled role. The defense: rewrite untrusted input into a neutral...
The PlayCoder benchmark is the cold shower the vibe coding movement needed. Researchers tested 10 state-of-the-art code LLMs on generating GUI applications across six categories. The models achieved high compilation rates. The code built and ran. But when they measured whether...
arXiv 2607.02857 identifies a vulnerability class where CLI commands that are each harmless alone form dangerous state relationships when an agent composes them. Across five real coding agents and five backend LLMs over 2,525 trials, they report 96.59% attack success under ben...
InceptionRAG fragments the payload into a chain of dormant passages, each benign under isolated inspection, that lead the model to self-deduce the target misinformation through multi-hop reasoning when retrieved together. Across three datasets and three LLMs it exceeds 80% ASR...
CS-Guard evaluated 9 guardrails across seven LLMs with 1,000 malware prompts, 7 jailbreaks and 331 code-to-code prompts covering infilling, completion and translation. Post-jailbreak text-to-code attack success averaged around 50%; code-to-code approached 100% on base models a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.