Research
BlueLM-GUI Trains a 35B-A3B Mobile Agent on Hundreds of Real Phones and Hits 84.9 on AndroidWorld
The report targets three deployment gaps in mobile GUI agents: sandbox training mismatching production, expensive real-device failures going unused, and fixed benchmarks saturating. Its answer is a dual-track pipeline with triple-system consensus evaluation plus an error correction and derivation module that salvages every trajectory into supervision, a three-stage recipe ending in agentic RL run on hundreds of real phones, and a quota-driven benchmark methodology upgradeable as the model improves. It reaches 87.4 on MobileGUI-VBench, 5.1 points above the best closed-source model, and 84.9 on AndroidWorld, the best open-source result reported.
↳ Follow the thread