发表机构
Harbin Institute of Technology, SCIR Lab; Peking University; Huawei Technologies Co., Ltd; Tsinghua University(哈尔滨工业大学; 北京大学; 华为技术有限公司; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EnvCraft通过自动化框架合成可执行环境和可扩展训练数据,解决智能体强化学习中交互环境稀缺问题,在类爪基准上提升11.9%。
AI 中文摘要
大语言模型的范式已迅速从被动的语言界面转向自主的类爪智能体,这些智能体能够在有状态的工作空间中执行长期任务。尽管智能体强化学习(Agentic RL)为优化这些智能体提供了一条有前景的路径,但其扩展严重受限于交互式训练环境的极度匮乏。现有的合成环境严格局限于工具调用端点,不足以满足类爪智能体端到端的现实需求。为弥补这一差距,我们提出了EnvCraft,一个用于合成可执行环境和可扩展训练数据的自动化框架。具体而言,EnvCraft采用环境合成引擎构建沙箱隔离的工作空间,并辅以拓扑感知的数据生成引擎生成连贯的任务轨迹。总体而言,我们合成了139个交互式环境,包含约2万个复杂任务,用于智能体强化学习训练。在Qwen3/3.5模型(8B-32B)上的实验表明,我们的方法在类爪基准上取得了高达+11.9%的提升,在通用工具使用基准上取得了+8.0%的提升,同时降低了推理令牌成本。结果证实,合成的可执行环境为训练提供了稳健且可泛化的学习信号。
英文摘要
The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its scaling is heavily bottlenecked by the severe scarcity of interactive training environments. Existing synthetic environments are strictly limited to tool-calling endpoints, rendering them insufficient for accommodating the end-to-end real-world demands of claw-like agents. To bridge this gap, we introduce EnvCraft, an automated framework for synthesizing executable environments and scalable training data. Specifically, EnvCraft employs an environment synthesis engine to build sandbox-isolated workspaces, alongside a topology-aware data generation engine to produce coherent task trajectories. Overall, we synthesize 139 interactive environments comprising approximately 20K complex tasks for Agentic RL training. Experiments on Qwen3/3.5 models (8B-32B) show that our method yields gains of up to +11.9% on Claw-style benchmarks and +8.0% on general tool-use benchmarks, with concurrent reductions in inference token cost. The results confirm that synthesized executable environments provide robust and generalizable learning signals for training.
Comments29 pages, 12 figures