arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14239cs.LG

CoArena:实时评估计算机使用与多智能体系统

CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time

Nitish Kovuru, Prateek Jannu

首次发表
浏览论文内容

中文总结 AI 辅助

CoArena提出实时评估计算机使用与多智能体系统的方法,通过用户实时提交任务、并发执行与在线评分,形式化定义实时性并给出完整评分方法,实现动态、抗污染的评估。

中文摘要 AI 辅助

计算机使用智能体的静态基准在发布时固定任务集,并针对该任务集对每个系统进行一次评分。这使得基准具有可复现性,但也使其偏离了本应衡量的目标:固定的任务集会老化、泄漏到训练语料库中,并且无法追踪人们逐周实际使用智能体的方式。CoArena直接衡量使用情况。真实用户提交任务;两个系统(每个系统是单个模型或位于相同工具接口之后的多智能体流水线)在相同的沙盒桌面环境中并发执行相同任务;用户在不了解结果由哪个系统产生的情况下对两个结果进行评判;公共排行榜根据这些评判重新拟合。核心贡献是对此类评估为何是实时的形式化描述。我们将实时定义为五个可测量的属性,每个属性都有公式和实例:连续任务到达、实时并发执行、在线评分更新、具有抗污染能力的新鲜度,以及从失败运行到可复用训练环境的有界反馈延迟。评分方法完整如下:Bradley-Terry成对模型、其带加权观测和平局的似然、惩罚最大似然估计器,以及单票到达时的流式更新(对同一似然执行随机梯度步骤,恢复Elo评分)。它通过观测信息和聚类稳健三明治给出置信区间,通过参数自助法给出排名带,给出新系统进入排行榜的规则,以及估计的收敛速率。投票质量通过评判者间一致性统计、冗余评判以及对平局和弃权(不执行)的显式处理来保证。一个包含211票的五系统示例从投票矩阵到评分、区间和排名带完整呈现。每个数字均来自既定输入或标记为说明性;没有一个是已部署系统的测量值。

英文摘要

Static benchmarks for computer-use agents fix a task set at release and score every system against it once. That makes them reproducible, and it lets them drift from what they should measure: a fixed task set ages, leaks into training corpora, and cannot follow how people actually use agents from week to week. CoArena measures use directly. Real users submit tasks; two systems, each a single model or a multi-agent pipeline behind the same tool interface, execute the same task concurrently in identical sandboxed desktops; users judge the two outcomes without knowing which system produced them; and a public leaderboard is refit from those judgments. The central contribution is a formal account of what makes such an evaluation real-time. We define real-time as five measurable properties, each with an equation and a worked example: continuous task arrival, live concurrent execution, online rating updates, freshness with contamination resistance, and bounded feedback latency from a failed run to a reusable training environment. The rating methodology follows in full: the Bradley-Terry pairwise model, its likelihood with weighted observations and ties, the penalized maximum-likelihood estimator, and the streaming update applied when a single vote arrives (a stochastic-gradient step on the same likelihood, recovering Elo). It gives confidence intervals from the observed information and a cluster-robust sandwich, rank bands from a parametric bootstrap, the rule by which a new system enters the board, and the convergence rate of the estimate. Vote quality is treated with inter-judge agreement statistics, redundant judging, and explicit handling of ties and abstentions. A five-system example with 211 votes is carried from the vote matrix to ratings, intervals, and rank bands. Every number is derived from stated inputs or labeled illustrative; none is a measurement of a deployed system.

发表机构

  • Coasty Research Lab(科斯蒂研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑