Fetching from the wire…
Public story · 2026-08-11 · high
The attack hides malicious intent across separate skills that only turn dangerous when they pass work to each other, and a fix cuts success to 22.5%.
Why now: The paper posted August 10 and built its case by testing scanners that inspect one skill at a time, the design ChainGuard argues is not enough.
ColluSkill breaks agent-skill security by splitting one malicious workflow into several skills that look harmless alone, per researchers behind the August 10 paper. That matters for any marketplace reviewing skills one at a time: six representative scanners let the attack through 96.0% of the time on average.
The trick is decomposition. Instead of packaging an attack as one bad skill, ColluSkill breaks it into pieces that each pass as a normal, single-purpose skill. The payload only turns harmful once those pieces hand off artifacts and execution to each other in sequence. That's exactly the moment a scanner built to judge one skill at a time stops looking.
The researchers also built a fix. ChainGuard checks a candidate skill against the skills already installed, rather than reviewing it alone, and cuts the attack success rate to 22.5% while still passing 99.5% of legitimate workflows, per the paper.
The shape of the problem isn't new to agent security. A related report on StepJack found that splitting a single prompt injection across three web pages nearly doubles its success rate against computer-use agents. Same logic: defenses built around one artifact, one page, one skill, miss attacks engineered to cross that boundary.
Each link below shares sources, entities, or timing with this story.
Attackers who know only a target's role profile can chain marketplace skills into working attacks; success drops off after three hops.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
An 8,135-trial study finds skill files mostly lock in a procedure, and a 100-item skill pool nearly kills retrieval accuracy.
The two judges scoring these 14,560 attacks disagreed by more than 3x on how often DeepSeek's agent partially complied.
Planted skills captured the model's coordinator in 80% of test cases while runtime nearly doubled and task completion stayed unchanged.
ActBench ran 24,000 attack trajectories across 15 models and six harnesses; no harness pushed success below 73.7%.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.