arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06197cs.AI

EnvACE:通过世界预演内化环境动态以实现智能体强化学习

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan Zhang, Xingshan Zeng, Weiwen Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出 EnvACE 方法,以世界预演替代外部环境交互训练 LLM 智能体,在多基准测试中性能优于基线,可突破外部环境约束扩展 LLM 智能体训练。

中文摘要 AI 辅助

针对长 horizon 工具使用的大型语言模型(LLM)智能体训练,通常依赖与真实或合成可执行环境的交互,这类环境的构建与验证成本高昂;或依赖难以落地的外部模拟器。本文提出 EnvACE,一种智能体强化学习方法,在训练过程中用世界预演替代与外部环境的交互。该策略在行动与预演间交替:首先生成工具调用,再扮演环境角色生成该动作对应的响应,并基于预演得到的响应调整后续决策。两个角色通过任务成功奖励进行端到端联合优化。通过世界预演,策略将动作与其环境响应的关系内化到自身参数中,形成可直接支撑决策的智能体世界模型。在 BFCL-v4、tau^2-Bench、VitaBench 和 FinMCP-Bench 基准测试中,EnvACE 表现出强劲且可迁移的性能,在整体评估中优于环境缩放基线。受控研究进一步表明,世界预演在各模型规模下均能持续提升策略学习效果。测试阶段,内化的世界模型允许在执行前进行私有预演,在适度的预演预算下无需额外外部交互即可进一步提升性能。本研究的发现确立了世界预演是突破外部环境约束、扩展 LLM 智能体训练的新路径。代码公开于此 https URL。

英文摘要

Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end-to-end using task-success rewards. Through world rehearsal, the policy internalizes the relationship between actions and their environment responses in its parameters, yielding an agent world model that directly supports decision making. Across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench, EnvACE achieves strong and transferable performance, outperforming environment-scaling baselines in the overall evaluation. Controlled studies further show that world rehearsal consistently improves policy learning across model scales. At test time, the internalized world model enables private rehearsal before committed execution, yielding further gains under a moderate rehearsal budget without additional external interaction. Our findings establish world rehearsal as a new path toward scaling LLM agent training beyond the constraints of external environments. Our code is publicly available at https://github.com/Within-yao/EnvACE.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Zhejiang University(浙江大学)
  • National University of Singapore(新加坡国立大学)
  • Sun Yat-sen University(中山大学)
  • Central South University(中南大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Tencent Inc.(腾讯公司)

机构由 AI 辅助整理,请以论文原文为准。

↑