Fetching from the wire…
Agents2026-08-02 · source-backed
Researchers reviewed public failure-scored trajectories from five web, enterprise-workflow, and desktop-control benchmarks: 10.7% were evaluator false negatives rejecting valid alternative solutions, 4.7% were broken or stale tasks. For the genuine failures, verification/feedback and planning errors dominate execution and grounding. Anyone comparing CUA products on leaderboard numbers is reading a figure with one in six negative results miscounted.
Each link below shares sources, entities, or timing with this story.
A June 26 paper (arXiv:2606.26294) describes a self-improving architecture where the agent and the evaluator that scores it evolve together, specifically to avoid the stagnation of optimizing against a fixed, gameable reward. (arXiv) Anyone building a self-improving harness ha...
RepoRescue tests coding agents on fixing a codebase broken by dependency and version changes across many files at once. That's a far more realistic and harder task than single-function bug fixing, and a better yardstick if you're deploying agents on legacy repos. Anyone who's...
arXiv 2608.24358 switched models mid-run on long coding tasks using cheap/expensive pairs from the Claude and GPT families. Full-trajectory escalation from weak to strong recovers under half the gap while costing a substantial premium, which the authors call the handoff tax. D...
2,910 programmatically verified tasks built from an ontology of 97 canonical UI components. Holding the harness fixed and changing only observation and action space, GPT-5 mini scores 83.1% with accessibility-tree observations and 48.9% with coordinate-only pixel control. Acro...
The PlayCoder benchmark is the cold shower the vibe coding movement needed. Researchers tested 10 state-of-the-art code LLMs on generating GUI applications across six categories. The models achieved high compilation rates. The code built and ran. But when they measured whether...
PAIChecker audits SWE-bench Verified and finds misalignment across five patterns and eleven scenarios, a direct consequence of the construction pipeline pairing a PR with whatever issue its description references, then using the issue as problem statement and the patch as test...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.