arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在具有挑战性的真实场景中对通用移动助手进行基准测试

Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

Yiqi Zhu, Feiyu Gao, Jiaxing Fan, Jiahui Zeng, Minggang Wu, Chenliang Li, Haiyang Xu, Peng Li, Ming Yan, Yang Liu

arXiv 2608.27477首次发表:更新:

发表机构

Tsinghua University; Alibaba Group; Institute for AI Industry Research (AIR), Tsinghua University(清华大学; 阿里巴巴集团; 清华大学人工智能产业研究院(AIR))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出名为GMA的移动助手基准,涵盖七个应用与300个不同难度任务,评估八个前沿模型发现其性能随任务复杂度提升而下降,还证实合适的智能体控制设计可显著提升性能。

AI 中文摘要

图形用户界面已成为评估自主AI智能体多模态交互任务的重要环境。现有基准如AndroidWorld和MobileWorld为移动智能体评估提供了坚实基础,但它们的应用覆盖范围和任务设计尚未完全反映现实移动使用的多样性与复杂性。我们提出GMA,一个用于在具有挑战性的真实场景中评估通用移动助手的基准。GMA基于开源项目引入了七个应用,涵盖生活方式分享、旅行规划等领域,以及四个难度等级的300个任务,从基础动作到复杂多步骤工作流不等。我们评估了八个前沿模型,发现随着任务复杂性增加,性能大幅下降,当前智能体仍远未可靠处理现实用户需求。我们还在共享环境、模型设置和任务分类下,对智能体控制(harness)选择进行了受控消融研究,包括上下文保留和显式状态跟踪。结果显示,合适的控制设计可显著提升性能,尤其在高要求工作流上,而特定设计的有效性会因基础模型而异。总体而言,GMA通过扩展应用覆盖范围和任务复杂性补充了现有基准,为评估移动智能体及研究控制设计如何支持复杂移动工作流中的可靠执行提供了具有挑战性的测试平台。

英文摘要

Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks. Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use. We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios. GMA introduces seven applications based on open-source projects, spanning domains such as lifestyle sharing and travel planning, and 300 tasks across four difficulty tiers, from atomic actions to complex multi-step workflows. We evaluate eight frontier models and find that performance declines substantially as task complexity increases, with current agents remaining far from reliably handling realistic user requirements. We further conduct controlled ablation studies of agent harness choices, including context retention and explicit state tracking, under a shared environment, model setting, and task taxonomy. Results show that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models. Overall, GMA complements existing benchmarks by expanding application coverage and task complexity, providing a challenging testbed for evaluating mobile agents and studying how harness design supports reliable execution in complex mobile workflows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑