Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications
arXiv·high signal
A study of 1,200 applications across six commercial LLM platforms found over 80% leak their system prompts under adversarial queries, traced to a mechanism the authors call 'attention drift.' Their AREA defense matches existing protection while improving usability by over 33%. A striking empirical result for anyone who assumes their system prompt is private.