arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38184cs.HC

RealGUINoise:真实界面噪声下GUI智能体鲁棒性的交互式跨平台基准

RealGUINoise: An Interactive Cross-Platform Benchmark for GUI Agent Robustness under Real-World Interface Noise

Yongjiang Wu, Junyuan Zhang, Ada Chen, Kuiyi Gao, Wenxuan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对真实界面噪声下GUI智能体性能与安全风险,提出交互式跨平台基准RealGUINoise,含42种噪声和7个框架,实验证明噪声降低性能并增加不安全行为。

中文摘要 AI 辅助

图形用户界面(GUI)智能体和计算机使用智能体(CUA)正迅速成为实用工具。然而,实际部署日益暴露出性能失败和安全风险,而这些问题的一个主要但尚未充分探索的来源在于日常界面的复杂且嘈杂的条件。现有基准大多假设干净环境或狭隘地聚焦于安全特定设置,并且缺乏一个统一的框架来跨不同智能体和平台进行一致、自动化的端到端评估。因此,我们引入了RealGUINoise,一个交互式跨平台、可扩展的基准,用于在完全交互式环境中系统评估GUI智能体在常见现实界面噪声下的表现。具体而言,RealGUINoise包含42种噪声类型,涵盖网页、桌面和移动任务,并集成了7个代表性智能体框架。它通过实时交互在真实世界日常任务上评估这些智能体,根据任务特定的黄金评分标准和干净环境轨迹,在可靠性、安全性和轨迹级行为方面比较其性能。我们的实验表明,这些噪声不仅降低任务性能,还显著改变智能体的行动轨迹并增加其不安全行为的倾向。这些发现暴露了干净环境中的能力与现实世界环境中可靠操作之间的关键差距,将RealGUINoise确立为开发更健壮和可信的GUI智能体的测试平台。

英文摘要

Graphical User Interface (GUI) agents and Computer-Using Agents (CUAs) are rapidly becoming practical tools. However, real-world deployment increasingly exposes performance failures and safety risks, while a major yet underexplored source of these problems lies in the complex and noisy conditions of everyday interfaces. Existing benchmarks largely assume clean environments or focus narrowly on security-specific settings, and lack a unified framework for consistent, automated end-to-end evaluation across diverse agents and platforms. Hence, we introduce RealGUINoise, an interactive cross-platform, extensible benchmark for systematically evaluating GUI Agents under common realistic interface noise in fully interactive environments. Specifically, RealGUINoise comprises 42 noise types spanning web, desktop, and mobile tasks and integrates 7 representative agent frameworks. It evaluates these agents on real-world daily tasks through real-time interaction, comparing their performance against task-specific golden rubrics and clean-environment trajectories in terms of reliability, safety, and trajectory-level behavior. Our experiments show that these noises not only degrade task performance but also substantially redirect agents' action trajectories and increase their propensity for unsafe behavior. These findings expose a critical gap between capability in clean environments and dependable operation in real-world settings, establishing RealGUINoise as a testbed for developing more robust and trustworthy GUI agents.

发表机构

  • Stanford University(斯坦福大学)
  • New York University(纽约大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑