Fetching from the wire…
Security2026-08-14 · source-backed
arXiv 2608.13010 scores top-five retrieval candidates against ranks 6–20 of the same query to spot answer-anchor concentration, and separately compares documents to lexically distinct neighbors to catch coordinated density before any query arrives. Deployed jointly, attack success falls from 67.4% to 14.0% while retaining 41.3% F1 on unpoisoned retrieval. Code at github.com/XrazyMee/RAGSieve. Using the corpus as its own reference is the part that makes this deployable.
Each link below shares sources, entities, or timing with this story.
Skill-α (arXiv 2608.01678) reframes skill generation as RL over sequential edits, decomposing skill construction into individually evaluable changes. The novel signal is a rollback reward that scores each modification by comparing downstream task execution using the original s...
Across 2,823 committed episodes on three frameworks, a one-class echo-state-network ensemble with CUSUM alarms catches 71% of mid-episode failures at a 5% false-alarm budget, three orders of magnitude cheaper than a judge call. But learned monitors don't transfer (AUROC 0.527...
Using op-schema-aware seeded fuzzing against a high-precision fp64 CPU reference on 24 Triton kernels, 15 correct and 9 intentionally buggy, the method caught all 9 buggy variants and passed all 15 controls across five GPU classes (arXiv:2606.20128). Standard kernel benchmarks...
PRISMA 2020 review, six databases, 743 records screened, 85 retained from 2023–2025 (arXiv 2608.10530). Perception-layer work (prompt injection, jailbreaking, adversarial perturbation) is 66% of papers. Action-layer vulnerabilities (tool misuse, code injection, sandbox escape)...
arXiv 2608.09624 separates harmful intent (a prompt property) from jailbreak success (an outcome from a specific model, decoder, and judge). On Llama, wrapping a prompt raises harmful generation from 0.05 to 0.27 while harmful-intent AUROC falls from 0.936 to 0.803. Attacks ge...
Tencent's SkillJack work (arXiv 2608.03509) shows self-evolving agents launder malicious intent during skill extraction. Attack success rates of 56.2% on SkillX and 89.2% on Anything2Skill, with 80% of implanted skills surviving deletion of the original poisoned records. If yo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.