APPSim-Bench:弥合真实应用与移动GUI智能体可复现评估之间的鸿沟
APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
浏览论文内容
中文总结 AI 辅助
提出APPSim-Bench,通过可控模拟应用兼顾真实性与可复现性,含557个任务,评估19个GUI智能体,最佳完成率仅50.27%,揭示移动端自主执行仍具挑战。
中文摘要 AI 辅助
移动图形用户界面(GUI)智能体能够根据自然语言指令执行任务,但其评估仍难以同时实现真实性与可复现性。现有基准通常在这两个目标之间进行权衡:简化应用缺乏真实世界的移动端复杂性,而真实商业应用则因推荐、广告、账户及动态内容引入不可控的变化。我们提出APPSim-Bench,通过可控的模拟应用来解决这一权衡问题,这些应用保留了与任务相关的交互逻辑,同时支持确定性评估。该基准通过编码智能体辅助且经人工验证的工作流程构建,包含覆盖17个高频中英文应用的557个任务。其可控的后端数据与基于结果的验证机制消除了环境随机性的主要来源,从而支持可复现的跨模型比较。在评估涵盖通用型与GUI专用型系统的19个GUI智能体时,我们发现自主移动端执行仍远未解决。最佳模型仅完成50.27%的任务,且28.55%的任务未被任何智能体解决。进一步分析表明,失败集中在较长工作流、数值推理任务以及以高动作开销和预算耗尽为特征的低效轨迹中。我们的项目可在以下网址获取:https://this-url。
英文摘要
Mobile GUI agents can execute tasks from natural-language instructions, but their evaluation remains difficult to make both realistic and reproducible. Existing benchmarks typically trade off these goals: simplified apps lack real-world mobile complexity, whereas live commercial apps introduce uncontrolled variation from recommendations, advertisements, accounts, and changing content. We propose AppSim-Bench, which addresses this trade-off through controllable simulated apps that preserve task-relevant interaction logic while supporting deterministic evaluation. Built through a coding-agent-assisted and human-verified workflow, it contains 557 tasks across 17 high-frequency Chinese and English apps. Its controllable backend data and outcome-based verification remove major sources of environmental stochasticity, enabling reproducible cross-model comparison. Evaluating 19 GUI agents, spanning general-purpose and GUI-specialized systems, we find that autonomous mobile execution remains far from solved. The best model completes only 50.27% of tasks, and 28.55% of tasks are not solved by any agent. Further analysis shows that failures concentrate in longer workflows, numerical reasoning tasks, and inefficient trajectories marked by high action overhead and budget exhaustion. Our project is available at https://github.com/Acrab-Agentic-Labs/AppSim.
发表机构
- Central China Normal University(华中师范大学)
- Acrab AI
- Agentic Labs
机构由 AI 辅助整理,请以论文原文为准。