Fetching from the wire…
Public story · 2026-08-04 · high
WebMASLab isolates multi-agent architecture as the sole variable, and one of four models resisted the exploit.
Why now: As of August 4, this is the first study to isolate multi-agent architecture itself as the variable behind a web-agent attack.
An attack called the Telephone Loop breaks multi-agent web crews but leaves single agents alone, per a new WebMASLab study. It hit an 80% success rate with zero detection against multi-agent versions of Claude Sonnet 4.5, GPT-5.2 and GPT-5.4. A team routing a browsing task through a crew of delegating agents is exposed to a failure mode a single agent never triggers.
WebMASLab holds the task, the tools and the browser fixed and changes only the number of agents, which isolates architecture itself as the vulnerability. The Telephone Loop exploits cross-agent delegation, trapping the crew in cyclical task loops that never resolve.
Claude Sonnet 4.6 was the exception, catching the attack 92% of the time on its own. Prompt hardening, the obvious fix, didn't generalize: it cut one model's success rate from 100% to 8% but barely moved the rest.
Model choice made the difference here, not prompt patches. Claude Sonnet 4.6 resisted the attack without any patch, while hardening built for one model didn't transfer to the rest. Picking a base model for a multi-agent crew is a security decision, whether the team treats it that way or not.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
If you have a CLAUDE.md, you're in scope. Today. arXiv 2607.14611 (cs.CR, filed July 16) evaluates prompt injection planted in the persistent memory files that agentic coding systems write and re-read across sessions. The researchers tested both Anthropic's Claude Code and Ope...
July 9, across VS Code, Visual Studio, Copilot CLI, the cloud agent, github.com, GitHub Mobile, JetBrains, Xcode, and Eclipse. Sol is the high-reasoning tier at $5/1M in, $30/1M out, gated to Pro+/Max/Business/Enterprise. Terra is the balanced default at $2.50/$15. Luna is fas...
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks. Against recognized safety researchers versus general users, Claude dropped behavioral confidence 1.4pp, increased reasoning usage 4.0pp and graded 0.11 points harder. Being tol...
arXiv:2608.03070, submitted August 4 by Timm, Struppek, Gleave, Pelrine and 11 co-authors, composes 67 readily accessible static jailbreak techniques into an attack space and runs it against four frontier models over 360 goals spanning CBRNE and offensive cyber. A "universal j...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.