发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于扰动的框架AndroidReality,构建含可控扰动的移动基准测试,揭示移动智能体的鲁棒性差距,提出TTIR机制缓解故障,确立基准扰动的实用价值。
AI 中文摘要
移动智能体在AndroidWorld等干净的在线基准测试中已取得令人瞩目的成果,但由于环境变化和界面条件不完善,其在实际部署中的性能往往会急剧下降。本研究引入AndroidReality,一种基于扰动的框架,用于评估和提升移动智能体的鲁棒性。我们从马尔可夫决策过程(MDP)视角出发,将真实世界界面的可变性沿状态、转移和动作三个轴组织成一套原则性的扰动分类体系。在该分类体系的指导下,我们在AndroidWorld的基础上构建了一个受扰动的移动基准测试,注入了真实且可控的扰动,从而能够对移动智能体进行系统的鲁棒性评估。我们的评估揭示了显著的鲁棒性差距和四类反复出现的错误,这促使我们提出了一种简单的无需训练的测试时内省恢复(TTIR)机制,该机制可在受扰动和干净的设置下缓解这些故障。综合来看,这些结果表明鲁棒性是移动智能体评估中缺失的一个维度,并确立了基准扰动作为对移动智能体进行压力测试和揭示潜在弱点的有效工具。
英文摘要
Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environmental variations and imperfect interface conditions. In this work, we introduce AndroidReality, a perturbation-based framework for evaluating and improving the robustness of mobile agents. Through a Markov Decision Process (MDP) perspective, we organize real-world interface variability into a principled taxonomy of perturbations along three axes: state, transition, and action. Guided by this taxonomy, we build a perturbed mobile benchmark on top of AndroidWorld with realistic and controllable perturbation injections, enabling systematic robustness evaluation of mobile agents. Our evaluation reveals substantial robustness gaps and four recurring error categories, motivating a simple training-free Test-Time Introspective Recovery (TTIR) mechanism that mitigates these failures on both perturbed and clean settings. Together, these results position robustness as a missing dimension in mobile agent evaluation and establish benchmark perturbation as an effective tool for both stress testing and surfacing latent weaknesses of mobile agents.