发表机构
Meituan; Fudan University; Shanghai Jiao Tong University; Zhejiang University(美团; 复旦大学; 上海交通大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多轮计算机使用智能体的在线强化学习问题,提出EvoCUA-1.5,通过STEPO、DTAC等方法应对多轮交互挑战,经实验验证其提高了训练稳定性和下游性能,为在线强化学习提供实用框架。
AI 中文摘要
计算机使用智能体必须通过与部分可观察、多模态桌面环境的重复交互来解决长期任务。尽管模仿学习和离线轨迹优化提供了强大的先验知识,但静态轨迹无法涵盖实际计算机使用的因果反馈循环。EvoCUA-1.5将自进化计算机使用智能体从离线经验学习扩展到在线强化学习,策略与可执行沙盒环境交互并从可验证任务结果中改进。在线强化学习在这种设置下需要更多操作。多轮交互引入了上下文管理观察、稀疏终端奖励、可变长度轨迹和缓慢环境反馈。EvoCUA-1.5通过步骤级策略优化(STEPO)、策略感知过滤和通过率校准、动态三自适应课程(DTAC)以及具有陈旧性控制和小组成批处理的完全异步强化学习基础设施来应对这些挑战。实验表明这些组件提高了训练稳定性和下游性能。EvoCUA-1.5在OSWorld-Verified上取得了63.2%的成功率,优于可比的32B/35B规模开放权重基线,甚至接近参数数量大得多的模型。总体而言,EvoCUA-1.5为扩展多轮计算机使用智能体的在线强化学习提供了一个实用框架。
英文摘要
Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static traces cannot cover the causal feedback loop of real computer use: each action changes the screen state, future action space, and recovery options. EvoCUA-1.5 extends self-evolving computer-use agents from offline experience learning to online reinforcement learning, where policies interact with executable sandbox environments and improve from verifiable task outcomes. Online RL in this setting requires more than directly reusing single-turn language-RL recipes. Multi-turn interaction introduces context-managed observations, sparse terminal rewards, variable-length trajectories, and slow environment feedback. EvoCUA-1.5 addresses these challenges with Step-Level Policy Optimization (STEPO), which preserves trajectory-level advantage balance after decomposition into step-level samples; policy-aware filtering and pass-rate calibration over verifiable synthesized tasks; Dynamic Tri-Adaptive Curriculum (DTAC), which combines learnable tasks, difficult positive replay, and controlled infeasible-task exposure; and a fully asynchronous RL infrastructure with staleness control and mini-group batching. Experiments show that these components improve training stability and downstream performance. EvoCUA-1.5 achieves 63.2\% success on OSWorld-Verified, outperforming comparable 32B/35B-scale open-weight baselines and even approaching models with significantly larger parameter counts. Overall, EvoCUA-1.5 provides a practical framework for scaling online RL in multi-turn computer-use agents.