arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CompoWorld:面向通用智能体的组合式环境扩展

CompoWorld: Compositional Environment Scaling for General Agents

Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan

arXiv 2609.33665首次发表:更新:

发表机构

AllSpark Team(全闪团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有环境生成局限于单一环境的问题,提出组合式环境扩展方法CompoWorld,通过组合服务库生成跨服务任务,结合SFT与RL训练,在八个基准上平均提升9.17分,并超越前沿模型。

AI 中文摘要

自动生成的环境为训练通用智能体提供了可扩展的交互数据来源。然而,现有方法主要是在单一环境内生成任务,而现实世界的工作流程要求智能体跨多个服务连接信息和动作。我们引入了组合式环境扩展(CompoWorld),它通过组合一个有限的可复用服务库来扩展任务空间。编码智能体将工具规范转化为具有类型化状态和共享接口的已验证服务,而世界模型则处理那些无法可靠实现的工具。一个随机游走过程通过依赖图连接服务,从而能够生成和验证需要信息跨服务流动的任务。验证过的轨迹支持监督微调(SFT),而我们的完成聚焦评分奖励(Completion-Focused Rubric Reward)通过强调每个回滚组内通过率较低的标准,引导强化学习(RL)朝向完整任务完成。我们构建了448个服务,暴露了10,130个工具,并使用3K条SFT轨迹和1K个RL任务来训练Qwen3.6-35B-A3B。实验结果表明,CompoWorld在八个基准测试中平均比其骨干模型提高了9.17分。在AutomationBench上,它超越了Claude Opus 4.6等前沿模型,并领先于所有参与比较的智能体专用35B-A3B模型。

英文摘要

Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (\textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑