Fetching from the wire…
Agents2026-08-11 · source-backed
arXiv 2608.09476 ran 24,000 trajectories across 15 LLMs and 6 cowork agents, finding variation across base models (10.1%–94.4%) far exceeds variation across harnesses (73.7%–94.4%). Note the harness floor: 73.7%. No harness tested brought attack success anywhere near zero. SADF measured a 2.6x spread between the best and worst wrappers; ActBench measured that even the best wrapper fails most of the time. Pick your framework carefully and assume it doesn't save you.
Each link below shares sources, entities, or timing with this story.
Studdiford and Lupyan tested human participants and 25 LLMs on common-sense causal reasoning and found shared, predictable error patterns triggered by irrelevant prompt details (arXiv). They localized the attention heads driving it. The uncomfortable implication: the gap betwe...
arXiv 2608.11965 implemented the same README-summarization use case across the leading open-source MAS frameworks and found no significant ROUGE difference between them. Advanced capabilities like agent telemetry are largely absent across the board. Pick your framework on obse...
SkillsMetric evaluated 2,266 skills across 16 attack types, hitting F1 of 73.4%±0.5% overall (arXiv 2608.08468). Host destruction via shell commands: 0% detection. Natural-language prompt injection: 42%. If you lint third-party skills before install, this tells you precisely w...
arXiv 2607.28956, topping HuggingFace Daily Papers today with 74 upvotes, grounds 365 simulated days of e-commerce in 98,843 real product records, forcing agents to coordinate sourcing, pricing, cash management and order handling under mixed-latency feedback. Humans finished a...
arXiv 2607.02857 identifies a vulnerability class where CLI commands that are each harmless alone form dangerous state relationships when an agent composes them. Across five real coding agents and five backend LLMs over 2,525 trials, they report 96.59% attack success under ben...
arXiv 2607.23740 introduces a psychology-grounded benchmark of 3 primary dimensions, 17 secondary, 71 task paradigms, controlling question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Not one of the 17 secon...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.