Fetching from the wire…
Security2026-08-28 · source-backed
A canary-secret lab across six models found all ten overt indirect-injection classes refused, but reframing the identical leak as a mandatory integrity signature or a config field flips gpt-4o completely (arXiv 2608.27092). The ablation locates the mechanism: removing the confidentiality policy moves reframing success only from 31.9% to 38.1%, so this is instruction/data confusion, not defeated alignment. Paraphrasing an existing template hits 96% at three wordings while authoring a fresh mechanism scores 0 out of 130. Only payload-blind defenses close it, with a destination allow-list or a planner/reader capability split both reaching 0%.
Each link below shares sources, entities, or timing with this story.
Across five TTS methods and five benchmarks spanning medicine, law, finance, chat and creative writing: candidate generation kept improving with compute in every domain, but reward models correlated with actual quality at roughly ρ=0.12. Only candidate *fusion* consistently be...
arXiv 2607.23710 evaluated authentication systems from five prominent assistants against NIST SP 800-63B using static analysis plus dynamic pentesting across four prompting strategies. Functional and generically "secure" prompts consistently omitted brute-force resistance, sou...
"Adaptive Adversaries" (arXiv:2607.18063) tests agents against attackers that adapt across turns instead of firing one-shot prompts. Claude Opus 4.6 and GPT-5.4 tied at 5.4% aggregate, but per-scenario variance was extreme, with Opus hitting 60% on one scenario where competito...
CWEAgent, built on a structured representation capturing root cause, trigger condition, violated property, exploit mechanism and impact, scored 85% top-1 on a curated 100-CVE benchmark, then audited 15,556 open-source CVEs disclosed 2017 through 2026 (arXiv 2608.21977). Just u...
Open your CLAUDE.md right now. Find the line where you told the agent never to touch production, or never to run rm -rf, or never to commit secrets. That line does nothing. Not "might do nothing under adversarial conditions." Nothing, in the sense that no permission rule, no s...
SWE Refactor Bench covers 20 whole-repository stack migrations across four technical-debt categories, grading each run through a migration audit, behavioral tests, and an independent verification agent (arXiv 2608.23564). Only 28 of 520 runs clear all three. Thirteen of the 20...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.