arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21557cs.AIcs.CL

OpenForgeRL:在任何环境中训练原生利用工具的智能体

OpenForgeRL: Train Harness-native Agents in Any Environment

Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现代AI智能体依赖复杂推理工具难端到端训练的问题,提出OpenForgeRL框架,通过轻量级代理和Kubernetes编排器,在多环境下对基于工具的智能体端到端训练,验证了框架效果并分析了工具选择和RL对智能体行为的影响。

中文摘要 AI 辅助

现代人工智能智能体依赖复杂的推理工具(如Claude Code、Codex和OpenClaw)来驱动多轮推理、工具使用和访问外部系统。这些强大但复杂的工具使得智能体难以通过开放基础设施进行端到端训练,因为其SFT/RL堆栈无法原生表达有状态的多进程工具推理。为此,我们提出了OpenForgeRL,一个用于在各种环境中对基于工具的智能体进行端到端训练的开源框架。OpenForgeRL通过一个轻量级代理实现这一目标,该代理在记录工具模型调用作为标准RL代码库(如veRL)的训练数据时为其提供服务,以及一个Kubernetes编排器,它在自己的远程容器中运行每次展开,共同实现了在任何环境中对任何工具进行大规模训练。通过解耦训练和推理,OpenForgeRL允许研究人员在智能体所部署的实际工具和环境中轻松地训练、研究和改进智能体。我们在各种复杂的工具和环境中验证了我们的框架,涵盖基于工具/爪子的智能体以及多模态GUI浏览器和计算机使用智能体。仅使用数百到数千个任务,OpenForgeClaw在ClawEval上达到31.7 pass^3和55.9 pass@3,在QwenClawBench上达到33.7。OpenForgeGUI在OSWorld-Verified上达到37.7,在Online-Mind2Web上达到63.0,在WebVoyager上达到72.3。两者在几乎所有基准测试中都优于类似规模的开放基线,并且在GUI设置中与几倍大的模型相匹配或超越。除了基准测试,我们分析了工具选择(如ZeroClaw、OpenClaw、Codex)和强化学习如何塑造智能体行为。我们发现一些工具比其他工具更难学习,并且强化学习提高了智能体的可靠性,如自我验证、工具覆盖和完成多步计划,尽管诸如错误恢复等关键能力仍然较弱。

英文摘要

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the real harnesses and environments they are deployed with. We validate our framework across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents. Using only hundreds to a few thousand tasks, OpenForgeClaw reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval and 33.7 on QwenClawBench. OpenForgeGUI reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager. Both outperform open baselines of similar size on nearly all benchmarks, and in the GUI setting match or surpass models several times larger. Beyond benchmarks, we analyze how harness choice (e.g., ZeroClaw, OpenClaw, Codex) and RL shape agent behavior. We find that some harnesses are substantially harder to learn than others, and that RL improves agentic reliability, such as self-verification, tool coverage, and completing multi-step plans, though critical abilities such as error recovery remain weak.

发表机构

  • Columbia University(哥伦比亚大学)
  • Dartmouth College(达特茅斯学院)
  • Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑