发表机构
The University of Hong Kong; Nanjing University; University of Science and Technology of China; National University of Singapore; Fudan University(香港大学; 南京大学; 中国科学技术大学; 新加坡国立大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出OSReward基准评估VLM评判者的可靠性,构建OS-Shepherd开源奖励模型缩小成本差距,为规模化可靠CUA奖励设计提供参考
AI 中文摘要
计算机使用智能体(Computer-using agents, CUAs)在数字领域发展迅速,其轨迹记录了智能体的动作、状态与推理过程。验证CUAs是否完成任务指令是CUAs评估、数据整理与强化学习的核心环节。人工编写的验证器或人工标注者无法规模化提供此类验证,因此该领域越来越多地采用视觉语言模型(Vision-Language Models, VLMs)作为CUAs轨迹的评判者,但一个长期未被研究的根本问题是:这些VLM评判者是否足够可靠?为系统研究该问题,我们推出OSReward,这是一个用于评估VLM对CUAs轨迹评判能力的真实高质量基准。这些轨迹来自不同智能体骨干在跨平台执行经人工验证的指令,再通过多阶段人工标注严格标注出真实判决结果。基于此,我们衍生出专注于真正困难案例的挑战集OSReward-Hard,以及用于细粒度效率与对齐评分的OSReward-Multi。对VLM评判者的迄今为止最全面评估发现,即使是最先进的模型也达不到理想评判者的标准,存在系统性宽松偏差,会错误将失败运行标记为成功。少数足够可靠可信任的模型规模化运行成本过高,而可负担的开源模型则远远落后。为缩小这一差距,我们构建并发布了OS-Shepherd-100K,这是一个面向CUA社区的带推理标注的轨迹判断开源语料库。我们在该语料库上训练了OS-Shepherd(9B和35B),这是提供低成本、稳定且可靠奖励信号的开源奖励模型,其成本比前沿商业评判者低30%-60%,且性能可媲美商业评判者。广泛分析进一步为规模化可靠CUA奖励的设计提供了参考。我们的代码、基准、数据集和模型检查点可在该https URL获取。
英文摘要
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, and are then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60x lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.
CommentsWork in progress