arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16798cs.CLcs.AIcs.LG

ClawGym II:探索智能体控制框架下的黑盒强化学习

ClawGym II: Exploring Black-Box RL on Agent Harness

Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao,… 展开作者

Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出统一黑盒强化学习框架,通过沙箱基础设施、前缀树轨迹重建等技术,实现通用智能体的稳定可扩展优化,在ClawGym-Bench等任务上取得性能提升。

中文摘要 AI 辅助

智能体控制框架通过协调智能体与环境的交互,大幅提升了长程任务的性能。然而,通过复杂控制框架进行强化学习的研究仍未得到充分探索,因为将此类训练扩展到长程智能体任务会带来根本性挑战。本研究提出了一种统一的黑盒强化学习框架,用于通过复杂控制框架实现通用智能体的稳定且可扩展的优化。具体而言,我们首先构建了基于沙箱的执行基础设施,该基础设施将任务环境和控制框架隔离在临时沙箱中,以支持大规模并发 rollout。随后,我们将策略优化与不透明的控制框架执行解耦,并在模型边界处放置一个服务代理以捕获模型调用。为了重建多轮轨迹并提高训练效率,我们将捕获的调用组织成前缀树,并进一步适配基于评论者的PPO和无评论者的GRPO,以在恢复的树结构上进行优化,同时在整个优化过程中保持训练-推理一致性。最后,我们引入了混合控制框架训练,允许单个模型由异构控制框架共同优化。使用Qwen3-30A3B,黑盒强化学习通过OpenClaw和Claude Code分别将ClawGym-Bench的Pass@1提升了9.98和14.81个百分点,同时在200-400个优化步骤内保持稳定。此外,该框架在更具挑战性的任务(如JobBench和OfficeQA)上也取得了一致的提升。总体而言,我们的框架支持通过黑盒控制框架实现通用智能体的有效、稳定且可扩展的优化,为异构执行系统提供了统一的训练支持。

英文摘要

Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses. Concretely, we first build a sandbox-based execution infrastructure that isolates task environments and harnesses within temporary sandboxes for large-scale concurrent rollouts. We then decouple policy optimization from opaque harness execution and place a serving proxy at the model boundary to capture model calls. To reconstruct multi-turn trajectories and improve training efficiency, we organize the captured calls into prefix trees and further adapt both critic-based PPO and critic-free GRPO to optimize over the recovered tree structure. Meanwhile, we maintain training-inference consistency throughout the optimization process. Finally, we introduce mix-harness training, allowing a single model to be jointly optimized by heterogeneous harnesses. With Qwen3-30A3B, black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectively, while remaining stable over 200-400 optimization steps. Moreover, the framework yields consistent gains on more challenging tasks such as JobBench and OfficeQA. Overall, our framework enables effective, stable, and scalable optimization of general agents through black-box harnesses, supporting unified training across heterogeneous execution systems.

发表机构

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
  • IQuest Research(智臻研究)

机构由 AI 辅助整理,请以论文原文为准。

↑