APPSim-Bench puts 19 GUI agents on 557 reproducible mobile tasks and the best one finishes half
arXiv·medium signal
arXiv 2609.07712 (submitted 2026-09-07) replaces live-app evaluation with 557 controllable simulated mobile app tasks so runs are deterministic and comparable. Across 19 agents the top performer completes 50.27% of tasks, which the authors read as autonomous mobile execution being far from solved. The reproducibility framing is the transferable part: the same argument applies to any agent benchmark currently run against a live third-party surface.