arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40284cs.LGcs.AIcs.CL

cua-speedrun:计算机使用智能体速度的标准化基准测试

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

发表机构卡内基梅隆大学
查看机构详情
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh

首次发表
浏览论文内容

中文总结 AI 辅助

提出cua-speedrun,一个标准化基准测试框架,用于评估计算机使用智能体的速度与效率,通过统一基础设施和任务集,发现推理努力、环境延迟等因素对性能、速度和成本的影响,并支持高效基准测试。

中文摘要 AI 辅助

计算机使用智能体(CUAs)通过图形用户界面(GUIs)在计算机上完成任务,最近在众多标准基准测试中已超越人类表现,包括困难的长时域任务。它们的能力无疑令人印象深刻,然而,阻碍CUAs广泛采用和部署的一个关键障碍仍然是其速度和成本。朝着更快且能力更强的CUAs的进展需要对其速度进行可靠评估,但许多CUA基准测试目前面临可复现性危机。基准测试基于复杂的基础设施,具有不同的机器和容器配置,这些配置混淆了对CUAs执行速度的评估。为了解决这一差距,我们提出了cua-speedrun,它引入了标准化的基础设施和任务集,专注于评估CUAs的速度和效率。cua-speedrun使用统一的虚拟机设置和执行流水线,以及一个通用智能体接口,使单一智能体实现能够无缝地在不同基准测试中运行。在四个不同的CUA基准测试中,我们评估了推理努力、智能体框架和环境延迟如何影响性能、速度和成本。我们发现没有一个模型系列在所有三个方面都是最优的;没有开放权重模型处于前沿,而且,出乎意料的是,对于某些模型,增加推理努力可以加速任务完成,而更快环境输入输出可能减慢整体任务完成时间。我们还证明,我们可以有效地减少大多数CUA基准测试的评估任务集,而不降低整体统计功效,从而实现更高效的基准测试和比较。我们相信cua-speedrun将推动朝着快速、高效的CUAs的结构化进展,解锁新的现实世界用例和应用。所有代码、基础设施和分析均可在此https URL获取。

英文摘要

Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. Across four different CUA benchmarks, we evaluate how reasoning effort, agent harnesses, and environment latency affect performance, speed, and cost. We find no single model family is optimal for all three; none of the open-weight models are on the frontier, and also, unintuitively, for some models increasing the reasoning effort can speed up task completion, while faster environment input-output can slow down overall task completion time. We also demonstrate that we can effectively reduce the evaluation task set of most CUA benchmarks without degrading overall statistical power, allowing for more efficient benchmarking and comparison. We believe cua-speedrun will enable structured progress towards fast, efficient CUAs, unlocking new real-world use cases and applications. All code, infrastructure, and analysis are available at https://cuaspeedrun.com.

↑