SCALECUA:通过可验证任务合成和高效在线强化学习扩展计算机使用代理
SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对计算机使用代理扩展能力受限问题,提出ScaleCUA框架,通过可验证任务合成及高效训练实现。包括设计VeriGen生成任务、前沿采样提高效率、视觉上下文分割加速训练,在相关平台取得新的最优性能。
AI中文摘要:
计算机使用代理(CUAs)正成为通过视觉感知和GUI执行自动化复杂数字工作流程的强大接口。具有可验证奖励的在线强化学习(RLVR)是扩展其能力的关键方向,但受可验证数据稀缺和在线强化学习效率低下的限制。为此引入ScaleCUA统一框架,通过可验证任务合成和高效训练扩展CUAs的在线强化学习。在数据层面,设计VeriGen通过迭代docker交互和多智能体反馈回路生成可验证强化学习任务,经共享docker交互探针扩展到100多个并发智能体工作者,产生24K +可验证任务和近3K高质量强化学习任务。提出前沿采样最大化样本效率,在训练方面设计视觉上下文分割平衡展开和训练引擎压力,使训练速度提高2.83倍。ScaleCUA在OSWorld上达到68.7%,在ScienceBoard上达到54.0%,在开源计算机使用代理中建立了新的最先进性能。
英文摘要:
Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scaling their capabilities. However, this paradigm is bottlenecked by verifiable data scarcity and online RL inefficiency. To break these barriers, we introduce ScaleCUA, a unified framework that scales online RL for CUAs via verifiable task synthesis and efficient training. At the data level, we design VeriGen, an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop. Scaled to 100+ concurrent agent workers via a shared docker interaction probe, this pipeline produces 24K+ verifiable tasks and nearly 3K high-quality RL tasks. To maximize sample efficiency, we propose Frontier Sampling, which tracks per-task capability and allocates rollouts to the current learning frontier. On the training side, we further design Visual Context Segmentation, a sliding window over recent visual context that balances rollout and training-engine pressure, yielding a 2.83x training speedup over step-wise decomposition. Together, ScaleCUA achieves 68.7% on OSWorld and 54.0% on ScienceBoard, establishing new state-of-the-art performance among open-source computer use agents. Code, models, and datasets are available at https://github.com/THUDM/SCALE-CUA.