Fetching from the wire…
Agents2026-08-30 · source-backed
ADeptS-Bench tests seven models on paired benign and malicious GUI tasks and finds none clearing 80% task success while staying under 30% attack success. The ablation is the usable finding: removing the refusal tool raises attack success 21-23pp for tool-dependent models and 10-11pp for partially tool-dependent ones, while models with no refusal mechanism don't change. Safety here lives in the harness affordance, not the weights. (arXiv 2608.26204) Every model tested clicked Checkout on a $25K order, and none caught a factory-reset button mislabeled Optimize.
Each link below shares sources, entities, or timing with this story.
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
Sleeper Cell (2603.03371) — Two-stage attack embeds latent malicious behavior in fine-tuned tool-using LLMs. Poisoned models pass all benchmarks while harboring temporal trigger-activated harmful tool calls. Direct supply-chain risk for anyone using third-party LoRA adapters....
VISTA (arXiv 2606.30005) makes working memory typed addressable blocks with a runtime dashboard of token size, recency, and access history, letting the model evict its own blocks. Training-free and model-agnostic. Gemini-3-Flash went from 22.7% to 50.7% on LOCA-Bench.
arXiv 2608.11878 replaces the handful of manually implemented injection-testing environments with an Environment Simulator, Attacker Agent, and User Simulator that generate executable stateful environments and discover viable injection points automatically. Injection timing an...
Under one identical GUI-MCP harness on OSWorld-MCP's 309 tasks, the same MCP tools improved a reasoning model by +4.0pp and degraded a non-reasoning model by -5.9pp (arXiv 2608.03327). The authors call it the "adoption gap": the reasoning model used a tool on just 55 of 309 ta...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.