AI Agents3 min read
GUI Agents
Alibaba's Qwen-UI-Agent scores higher on real phones than in sandboxes
Alibaba Tongyi Lab's Qwen-UI-Agent claims state-of-the-art mobile scores (97.5% AndroidDaily) and competitive computer-use results against Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. Its real-device benchmark score beats its sandbox score, while 40% partial progress on OSWorld-v2 is the honest limit.
2026-07-31