OSWorld-Pro:面向计算机使用智能体的基于过程的评估
OSWorld-Pro: Process-based Evaluation for Computer Use Agents
浏览论文内容
中文总结 AI 辅助
本文提出OSWorld-Pro,一个包含300多个任务和2800多个子目标、基于67,000条人工标注的过程化评估基准,用于揭示计算机使用智能体在子目标层面的失败模式,并发现即使顶级模型Claude Opus 5也仅达75.7%的准确率。
中文摘要 AI 辅助
计算机使用智能体(CUA)的评估通常仅限于其(在数百步结束时)创建的最终交付物,并使用功能验证器进行评估,如OSWorld中所示。然而,这种对最终状态性能的评估缺乏对智能体在各种任务中如何以及为何失败的透明度,从而掩盖了后续改进的关键见解。例如,在键盘输入中出错的智能体所需的缓解策略,与在图形用户界面上未能精确提供基于点击的输入的智能体所需的策略不同。我们引入了OSWorld-Pro:一组包含超过300个任务、超过2800个子目标的数据集,以实现基于超过67,000条人工标注的CUA程序化评估。我们使用稳健的、与人类对齐的LLM评判器来评估OSWorld-Pro子目标的完成情况,从而揭示模型在一系列顺序依赖的子目标中的进展。我们的研究结果表明,即使对于最先进的LLM,OSWorld-Pro也具有挑战性,顶级模型如Claude Opus 5仅达到75.7%,而在OSWorld上为83.4%。此外,我们识别了各种模型的关键的、以过程为中心的失败模式(例如,与子目标无关的动作和基于点击的错误),以提供改进CUA性能和效率的见解。
英文摘要
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agents fail in various tasks, obfuscating critical insight for subsequent improvement. For instance, agents that err during keyboard inputs would require a different mitigation strategy from those that fail to precisely provide click-based inputs on the graphical UI. We introduce OSWorld-Pro: a set of over 300 tasks containing over 2800 subgoals to enable the procedural evaluation of CUAs grounded in over 67,000 human annotations. We use robust human-aligned LLM-Judges to evaluate the fulfillment of OSWorld-Pro subgoals and thereby reveal the progress that models make throughout a series of sequentially dependent subgoals. Our findings reveal that OSWorld-Pro is challenging even for state-of-the-art LLMs, with top performers like Claude Opus 5 achieving only 75.7% vs. 83.4% on OSWorld. Furthermore, we identify critical process-focused failure modes of various models (e.g. subgoal-irrelevant actions and click-based mistakes) to provide insights to improve performance and efficiency of CUAs.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。