arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

交互式奖励智能体:通过环境状态验证进行图形用户界面任务评估

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Chenrui Shi, Yuwei Wu, Yang Liu, Ruining Feng, Zirui Shang, Zhi Gao, Lifeng Fan, Che Sun

arXiv 2607.25904首次发表:更新:

AI 中文总结

研究图形用户界面任务评估难题,提出交互式奖励智能体IRA,基于提议验证框架,结合可见界面与环境状态证据。IRA在GUI-RewardBench基准测试中准确率达86.9%,应用于强化学习时实现34.0%的OSWorld成功率,能为训练图形用户界面智能体提供有效奖励信号。

AI 中文摘要

图形用户界面任务评估旨在确定图形用户界面智能体是否成功完成用户指令。自动图形用户界面任务评估受到越来越多关注,因其评估结果可作为测试时扩展和训练后奖励信号。然而,可靠的图形用户界面任务评估仍具挑战性,因为判断常需访问执行轨迹截图之外的环境状态。本文提出基于提议然后验证框架的交互式奖励智能体(IRA),从执行后环境获取并验证证据。给定任务指令和智能体执行后的图形用户界面环境,IRA先提出任务完成条件,再通过调用系统、应用和图形用户界面工具进行验证。此设计在交互过程中结合了可见界面和环境状态的证据。我们还引入了GUI-RewardBench,一个涵盖10个Ubuntu桌面应用类别的321个图形用户界面任务轨迹的基准。实验表明,IRA在GUI-RewardBench上准确率达86.9%,优于现有评估基线。我们还将IRA应用于图形用户界面智能体的强化学习,实现了34.0%的OSWorld成功率,证明IRA可为训练图形用户界面智能体提供有效奖励信号。

英文摘要

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑