发表机构
Zhejiang University; Yale University; University of Chinese Academy of Sciences; Tongji University(浙江大学; 耶鲁大学; 中国科学院大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对Android GUI智能体抗运行时异常鲁棒性不足的问题,提出AnTrap基准评估框架,发现16款领先GUI模型普遍受动态异常影响,且深度上下文陷阱暴露模型内在局限性。
AI 中文摘要
GUI智能体在Android设备上部署时经常遇到动态异常,从意外弹窗到动作误用,然而现有基准缺乏对智能体抗运行时异常鲁棒性的系统性评估。我们引入AnTrap,这是一个向智能体执行轨迹注入动态扰动的综合基准。我们提出一种分类法,将现实世界异常组织为四个层级(状态、思考、动作和回合),包含十个细分子类别,并开发了一种构建流水线,在引入现实对抗条件的同时保留任务可解性。通过评估16个领先的GUI模型,我们发现它们普遍易受动态异常影响,即使是最强的模型也会出现显著的性能下降。此外,我们在原始环境和对抗环境中进行GRPO训练以验证我们的基准,将环境可学习的异常与推理瓶颈导致的异常区分开来。我们的研究结果表明,虽然状态和动作层的单步陷阱在很大程度上可通过对抗强化学习解决,但深度上下文陷阱(如状态死锁)暴露出仅在含陷阱的环境中训练无法解决的内在局限性。
英文摘要
GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against runtime anomalies. We introduce AnTrap, a comprehensive benchmark that injects dynamic perturbations into agent execution trajectories. We propose a taxonomy organizing real-world anomalies into four layers (State, Thinking, Action and Round) with ten fine-grained subcategories, and develop a construction pipeline that preserves task solvability while introducing realistic adversarial conditions. Evaluating 16 leading GUI models, we reveal universal vulnerability to dynamic anomalies, with even the strongest models suffering significant performance degradation. Furthermore, we conduct GRPO training in both original and adversarial environments to validate our benchmark, separating environment-learnable anomalies from reasoning-bottlenecked ones. Our findings show that while single-step traps at state and action layers are largely addressable through adversarial reinforcement learning, deep contextual traps, like state deadlock, expose intrinsic limitations that cannot be resolved by training in environments with traps alone.