Fetching from the wire…
Public story · 2026-07-31 · high
It also hit 92.2% on MobileWorld-Real, the benchmark where Alibaba says it matches Opus 4.8 and Gemini 3.1 Pro.
Why now: The technical report has pushed Qwen-UI-Agent to number one trending on Hugging Face, with 242 upvotes as of July 31.
Alibaba's Tongyi Lab built a GUI agent that clears 97.5% of AndroidDaily tasks, per its technical report on arXiv. It also scored 79.5% on OSWorld-Verified, a harder desktop-focused suite. Qwen-UI-Agent is meant to automate screen work end to end: tapping through apps, filling forms, running command-line jobs, without a person watching every step.
On MobileWorld it scored 82.1%. On MobileWorld-Real, 92.2%. Tongyi says those numbers are competitive with Opus 4.8, Gemini 3.1 Pro and GPT-5.6 Sol on the same tests.
The agent uses one action space across mobile, desktop, web and a DeepSearch mode. It mixes GUI taps and clicks with CLI commands. Instead of issuing one action at a time, it batches several into a single model turn.
Tongyi trained it with online reinforcement learning on trajectories running past 100 turns. The training ran across more than 10,000 concurrent environments, per the report. That's a lot of parallel practice for a model whose job is clicking through other people's software correctly on the first try.
Scores this clean are worth a second look. AndroidDaily and OSWorld-Verified are curated test suites. A model tuned hard against them can still choke on the one app or layout that never showed up in training.
Each link below shares sources, entities, or timing with this story.
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
Alibaba released Qwen 3.5, a 397B MoE model (17B active per token) that can see and control desktop apps, mobile apps, and web browsers by processing UI screenshots and executing multi-step workflows autonomously. 60% cheaper to run than its predecessor, 8x higher throughput,...
arXiv 2607.29199 tests three frontier GUI agents under screen-grounded, user-side persuasion, with no environment injection at all. A single-line guardrail cuts attack success rate by ~40 points in single-turn scenarios. Four-turn escalation chains push guarded ASR back up by...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
A 2.4-trillion-parameter MoE multimodal model with 1M-token context, aimed at long-horizon autonomous software work. Alibaba reports 86.1 against 83.2 for GPT-5.6 Sol Max and 85.0 for Fable 5, the first credible claim that a Chinese lab leads on GUI-driving agentic benchmarks....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.