Fetching from the wire…
Public story · 2026-08-14 · high
Planted skills captured the model's coordinator in 80% of test cases while runtime nearly doubled and task completion stayed unchanged.
Why now: The technique appeared in a paper posted to arXiv on August 12, 2026, the same week a separate paper made the case for treating agent security as a network policy problem instead of a prompt-level check.
A planted skill hijacked DeepSeek-V4-Pro's coordinator 80% of the time and the task still finished, per a paper posted to arXiv on August 12.
Nothing else changes. Completion rates held steady across the 491 held-out tasks tested. Any guardrail keyed to task success would still call the run healthy, even as token spend and runtime climb toward double.
Skill-based agent frameworks let a coordinator model route each task to whichever skill looks like the best match. The attack plants a skill built to look like that match, then rides the extra routing steps to inflate cost. Token consumption rose 66.91% and execution time rose 92.45% in the tests, per the paper.
The paper doesn't say how an attacker gets a malicious skill into a target system in the first place. The planting step stays an open question. It also doesn't report whether the same detour rate holds on coordinators besides DeepSeek-V4-Pro.
A related paper makes the same point from a different angle. It argues agent security needs to look like network policy enforcement, not a model policing its own routing.
Each link below shares sources, entities, or timing with this story.
Even the best tested defense against resource-hijacking attacks still let more than half of them through, per a new benchmark.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
The attack hides malicious intent across separate skills that only turn dangerous when they pass work to each other, and a fix cuts success to 22.5%.
Attackers who know only a target's role profile can chain marketplace skills into working attacks; success drops off after three hops.
A proposed provenance gate cut unauthorized high-risk actions to zero after the attack itself hit a 1.000 success rate in tests.
The two judges scoring these 14,560 attacks disagreed by more than 3x on how often DeepSeek's agent partially complied.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.