arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11050cs.AIcs.LG

AgentHorizon:评估用于长周期计算机使用任务的智能体评判器

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

Xing Han Lù, Dheeraj Vattikonda, Sina Hajimiri, Fatemeh Pesaran Zadeh, Parishad BehnamGhader, Ghazwa Darwiche, Amirhossein Kazemnejad, Christopher Pal, Alexandre Drouin, Siva Reddy

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出AgentHorizon基准,评估11个智能体评判器在长周期计算机使用任务上的表现,发现GPT-5.5在AH子集平衡准确率达80.9%,工具使用对不同模型影响不同,需能验证长交互中任务完成情况的评判器。

中文摘要 AI 辅助

计算机使用智能体能够完成复杂任务,这使得自动评判器的应用日益广泛,无论是用于训练还是无需人工参与的评估。尽管自动评判器具有灵活性,但它们在跨越多个应用的长任务上的可靠性仍不明确。由长序列截图和操作组成的轨迹可能看似完整,但实际上可能违反指令约束或产生意外副作用。为识别这些错误,评判器需要结合用户指令对轨迹进行检查。为此,我们推出AgentHorizon,这是一个包含1373个计算机使用任务(指令-轨迹对)的基准,这些任务来自166小时的人类记录轨迹,涵盖三个操作系统。通过记录密切相关指令的轨迹,我们可以通过交换指令构建负任务,这种配对设计用于评估评判器区分成功轨迹与完成相似(但不兼容)请求的轨迹的能力。我们发布的基准分为三个子集:前沿子集AgentHorizon(AH)、简化子集AgentHorizon-Simple(AH-S)和开发子集AgentHorizon-Development(AH-D)。我们进一步评估了11个评判器,评估方式包括:(1)直接传入完整轨迹(最多300张截图和操作);(2)将它们作为编码智能体应用于五个智能体框架。我们发现,我们最优的智能体评判器GPT-5.5在AH子集上达到了80.9%的平衡准确率。我们还发现,工具使用提升了部分模型的性能,但对于开放权重模型却导致性能下降,且不同评判器在接受有效轨迹和拒绝失败轨迹的能力上存在巨大差异。我们的研究结果强调,需要能够在长交互历史中定位并验证任务是否正确完成的评判器。

英文摘要

Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of long sequences of screenshots and actions, may appear complete, but in reality violates constraints from the instruction or introduces an unwanted side effect. To identify these errors, a judge needs to examine the trajectory with respect to the user's instruction. To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems. By recording trajectories for closely related instructions, we can construct negative tasks by swapping the instructions. This paired design evaluates judges on their ability to distinguish a successful trajectory from one that completed a similar (but incompatible) request. We release the benchmark under three splits: a frontier split, AgentHorizon (AH), a simplified split, AgentHorizon-Simple (AH-S), and a development split, AgentHorizon-Development (AH-D). We further evaluate eleven judges by (1) directly passing the full trajectory (with up to 300 screenshots and actions), and (2) using them as coding agents across five agent harnesses. We find that our best agentic judge, GPT-5.5, achieves 80.9% balanced accuracy on the AH subset. We find that tool-use improves certain models but results in worse performance for open-weight models, and that judges differ drastically in their ability to accept a valid trajectory and reject failed ones. Our findings highlight the need for judges that are capable of locating and verifying often hidden evidence that a task was properly completed inside long interaction histories.

发表机构

  • ServiceNow Research(ServiceNow研究院)
  • McGill University(麦吉尔大学)
  • Mila – Quebec AI Institute(米拉-魁北克人工智能研究所)
  • ÉTS Montréal(蒙特利尔高等技术学院)
  • Seoul National University(首尔大学)
  • Université Laval(拉瓦尔大学)
  • Polytechnique Montréal(蒙特利尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑