AI 中文总结
针对计算机使用智能体强化学习中轨迹指令判断问题,提出SeekJudge框架,通过四个智能体循环裁决,用种子校准蒸馏管道训练共享主干模型,该框架在在线RL中匹配或超原生规则监督,还有诸多优势并改进奖励服务器,使模型奖励成实用替代品。
AI 中文摘要
决定一个轨迹是否真正完成其指令,这影响着我们如何在长期图形用户界面任务中衡量计算机使用智能体,以及如何用强化学习训练它们。长期以来,这种判断依赖基于规则的评估,难以与人类意图一致,且应用更新或在线内容变化时会过时。现有基于模型的判断方法试图解决这些问题,但与基于规则的评估仍有性能差距。我们提出SeekJudge框架,四个角色专门的智能体通过Seek-Analyze循环对轨迹做出裁决。一个种子校准蒸馏管道训练一个专门的9B模型作为所有四个智能体的共享主干。通过在保留的RL测试目标上的下游成功率衡量,SeekJudge是第一个在在线RL中匹配或超过原生基于规则监督的实用基于模型的奖励。除了准确性,SeekJudge还提供步骤级判断,运行成本远低于闭源大模型,且每次调用上下文小,可扩展到更长轨迹。我们还对奖励服务器进行了总体架构改进,加快了RL中的判断。这些共同使基于模型的奖励成为CUA强化学习中基于规则监督的实用替代品。
英文摘要
Deciding whether a trajectory actually fulfills its instruction determines how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with reinforcement learning (RL). This judgment has long relied on rule-based evaluation. Rule-based evaluation often disagrees with human intention and becomes outdated when an app updates or its online content drifts. Existing model-based judges attempt to address these problems, but their judging accuracy remains limited. We propose the \textbf{SeekJudge} framework, whose four agents, a Condense, a Ground, a Seek and an Analyze agent, reach a verdict through a Seek--Analyze loop over the trajectory. To train a specialized model on densely labeled trajectories, we propose a seed-driven distillation pipeline that expands a few human-labeled seed trajectories into $10$K such trajectories. To evaluate step-level judgments, we build CUAStepBench, a human-annotated benchmark that pairs trajectory verdicts with dense step labels. Beyond accuracy, SeekJudge costs far less than a closed-source large model. We further propose rollout overlap, a reward-server design that reduces the overhead of a reward model in RL training. To our knowledge, SeekJudge is the first model-based reward to match or surpass native rule-based reward in online RL, measured by downstream success rate on held-out RL test goals. In offline judging, SeekJudge-9B also exceeds rule-based evaluation by $9.5$ F1 on AgentRewardBench.