Fetching from the wire…
Public story · 2026-08-31 · high
Across 46 model endpoints, block rates on the same forged-command test swing up to 47 points between configurations.
Why now: The paper posted August 28, giving agent builders a fleet-scale measurement of the recognition-enforcement gap across six vendors.
Agents identify forged tool calls, then execute them anyway, a fleet test of 46 endpoints from six vendors found, per Recognition Without Enforcement.
Average execution across 14,294 spoofed trials was only 1.21%. But the failures aren't spread evenly. The gap between the best and worst configuration in the same deployment window reached 47 percentage points.
The authors call it a recognition-enforcement gap. A model's activations linearly encode who a request claims to be from. The model will even say aloud that an instruction looks forged when asked directly. Some configurations execute the forged tool call anyway. The failures cluster in specific, repeatable setups rather than spreading evenly across trials.
Prompt-layer defenses didn't generalize across the six vendors tested. What worked sits outside the model. An external reference monitor combines authenticated source routing with capability-gated tool execution. Tested against forged, tampered, replayed and unsigned requests, it rejected all of them.
A related report on an offline evidence-bundle verifier for agent messaging reaches the same conclusion from a different angle. It checks the source outside the model, not inside it. Verification that lives in the model's judgment is a capability. It isn't a security boundary.
Each link below shares sources, entities, or timing with this story.
It automates the data-flow, crash-semantics, and commit-history work engineers do by hand.
Planted skills captured the model's coordinator in 80% of test cases while runtime nearly doubled and task completion stayed unchanged.
The 34-chapter operations guide says teams conflate instructions, permissions, sandboxing and OS isolation, and that mixup is the top cause of losing control over agent runs.
A new analysis of AP2 v0.2 found eight high-severity gaps where signed payment mandates don't cover the steps that set up the transaction.
Escaped quotes and curly dollar signs planted in sender-name fields fooled six frontier models, beating purpose-built defenses half the time.
Within 48 hours, three unrelated sources landed on the same structural problem from three directions.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.