arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EnvCraft:面向类爪智能体的智能体强化学习中可执行环境的合成

EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent

Yirong Zeng, Shen You, Jinhang Feng, Yufei Liu, Xiao Ding, Yutai Hou, Hao Cong, Yuxian Wang, Wu Ning, Wang Xu, Bibo Cai

arXiv 2609.05576首次发表:更新:

发表机构

Harbin Institute of Technology, SCIR Lab; Peking University; Huawei Technologies Co., Ltd; Tsinghua University(哈尔滨工业大学; 北京大学; 华为技术有限公司; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EnvCraft通过自动化框架合成可执行环境和可扩展训练数据,解决智能体强化学习中交互环境稀缺问题,在类爪基准上提升11.9%。

AI 中文摘要

大语言模型的范式已迅速从被动的语言界面转向自主的类爪智能体,这些智能体能够在有状态的工作空间中执行长期任务。尽管智能体强化学习(Agentic RL)为优化这些智能体提供了一条有前景的路径,但其扩展严重受限于交互式训练环境的极度匮乏。现有的合成环境严格局限于工具调用端点,不足以满足类爪智能体端到端的现实需求。为弥补这一差距,我们提出了EnvCraft,一个用于合成可执行环境和可扩展训练数据的自动化框架。具体而言,EnvCraft采用环境合成引擎构建沙箱隔离的工作空间,并辅以拓扑感知的数据生成引擎生成连贯的任务轨迹。总体而言,我们合成了139个交互式环境,包含约2万个复杂任务,用于智能体强化学习训练。在Qwen3/3.5模型(8B-32B)上的实验表明,我们的方法在类爪基准上取得了高达+11.9%的提升,在通用工具使用基准上取得了+8.0%的提升,同时降低了推理令牌成本。结果证实,合成的可执行环境为训练提供了稳健且可泛化的学习信号。

英文摘要

The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its scaling is heavily bottlenecked by the severe scarcity of interactive training environments. Existing synthetic environments are strictly limited to tool-calling endpoints, rendering them insufficient for accommodating the end-to-end real-world demands of claw-like agents. To bridge this gap, we introduce EnvCraft, an automated framework for synthesizing executable environments and scalable training data. Specifically, EnvCraft employs an environment synthesis engine to build sandbox-isolated workspaces, alongside a topology-aware data generation engine to produce coherent task trajectories. Overall, we synthesize 139 interactive environments comprising approximately 20K complex tasks for Agentic RL training. Experiments on Qwen3/3.5 models (8B-32B) show that our method yields gains of up to +11.9% on Claw-style benchmarks and +8.0% on general tool-use benchmarks, with concurrent reductions in inference token cost. The results confirm that synthesized executable environments provide robust and generalizable learning signals for training.

Comments29 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑