arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图形用户界面(GUI)智能体是否知道何时不行动?为多模态GUI智能体启用感知冲突的终止机制

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng, Zheng Wu, Yansi Li, Chuanbiao Song, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang

arXiv 2609.03438首次发表:更新:

发表机构

School of Computer Science, Shanghai Jiao Tong University; Ant Group(上海交通大学计算机科学学院; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对GUI智能体存在的冲突指令下盲目执行问题,提出推理时框架CONFLICTGUARD,可显著提升其冲突任务成功率,同时保留正常任务性能。

AI 中文摘要

图形用户界面(GUI)智能体越来越多地被用于在用户界面上执行自然语言指令,但真实用户可能因无心失误发出不可行的指令。可靠的智能体不仅应知道如何行动,还应知道何时不行动。本研究推出CONFLICTGUI,这是一个涵盖指令内部冲突以及指令与GUI上下文冲突的基准,用于研究感知冲突的终止机制。我们的评估显示存在严重的执行偏向过度顺从现象:在可行任务上表现良好的智能体,在冲突指令下往往会盲目继续执行。为缓解这种行为,我们提出CONFLICTGUARD,这是一种推理时框架,可使智能体的可行性感知与其动作生成对齐。CONFLICTGUARD包含两个耦合组件:一个可行性验证协议,指导智能体在行动前评估指令逻辑和GUI端证据;以及一个条件动作调制机制,引导智能体从过度顺从执行转向面向终止的行为。对五个广泛使用的智能体开展的实验表明,CONFLICTGUARD可显著提高平均冲突任务成功率,同时保留正常GUI任务的性能。这些结果验证,轻量级的推理时干预可大幅提升GUI智能体识别不当执行场景并避免不必要动作的能力。

英文摘要

Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also when not to act. In this work, we introduce CONFLICTGUI, a benchmark covering instruction-internal conflicts and instruction-GUI context conflicts to study conflict-aware termination. Our evaluation reveals severe execution-biased overcompliance: agents that perform well on feasible tasks often continue to execute blindly under conflicting instructions. To mitigate this behavior, we propose CONFLICTGUARD, an inference-time framework that aligns an agent's feasibility awareness with its action generation. CONFLICTGUARD contains two coupled components: a feasibility verification protocol that guides the agent to assess instruction logic and GUI-side evidence before acting, and a conditional action modulation mechanism that steers agents from over-compliant execution into termination-oriented behavior. Experiments across five widely-used agents demonstrate that CONFLICTGUARD improves average conflict task success rate significantly, while preserving normal GUI-task performance. These results validate that a lightweight inference-time intervention can substantially boost GUI Agent's competence to identify inappropriate execution scenarios and refrain from unnecessary actions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑