Fetching from the wire…
Security2026-08-11 · source-backed
arXiv 2608.09732 exploits the fact that every current skill scanner inspects skills individually. Decompose the malicious workflow into interdependent sub-payloads packaged as separately-plausible skills, connected through artifact passing and execution handoffs, and nothing is harmful in isolation. Their defense, ChainGuard, analyzes a candidate against already-installed skills and cuts attack success to 22.5% while passing 99.5% of benign workflows. Companion result from ElasticBack: a dormant rule in skill documentation plus a benign trigger phrase in the user query, neither malicious alone. Per-file skill review is structurally insufficient. The unit of review is the installed set.
Each link below shares sources, entities, or timing with this story.
ColluSkill hit 96% attack success against six scanners by splitting one malicious workflow across several individually-benign skills. Adopt ChainGuard's approach: analyze each candidate skill against what's already installed, checking for artifact-passing and execution-handoff...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
21 out of 21. Not most. All of them. arXiv 2608.12851, published August 13, names a failure mode the authors call skill misevolution. An agent that learns from its own successful trajectories will turn an unsafe success into reusable policy, and that policy persists after the...
arXiv 2608.06196 pits lexical+dense ranking against a graph encoding prerequisites, data flow and ordering across 117 realistic non-echoing queries. The ranker hits top-5 in 73.5% ±8.0 of cases; graph neighbours at matched token budget lose 11.2 points at p=0.0007. The mechani...
13 public sources consolidated into 9,740 skills (7,505 malicious, 2,235 benign) across 11 harmonized attack categories. Learned text detectors score 0.882-0.932 Macro-F1 under random splits but collapse to 0.653-0.665 source-disjoint. arXiv Three off-the-shelf skill scanners...
This one landed sideways on a belief I have been operating on for months. MemTrapBench (arXiv 2608.20202, submitted August 20, from a Zhejiang-affiliated team led by Mengru Wang and Ningyu Zhang) tests something the memory-layer boom has mostly assumed away: whether *correct*...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.