发表机构
WeChat AI, Tencent(微信人工智能,腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体强化学习环境管理开销大的问题,提出WeEnv,通过层组打包、按需初始化及弹性资源供应,将初始化时间减少5.6-14.2倍,迭代时间占比降至9.1%,并已在微信部署。
AI 中文摘要
智能体强化学习(Agentic RL)与传统强化学习不同,每个任务都在复杂环境(如虚拟机或容器)中执行。我们发现智能体强化学习需付出高昂的环境代价:迭代时间中很大一部分用于环境而非学习。根本原因在于缺乏环境管理的全生命周期解决方案。我们提出WeEnv,它在打包、初始化和供应方面管理环境。WeEnv将组件打包为独立发布的层组,并在初始化时进行组合,从而更新一个组件只需重新发布一个小型层组,而非包含该组件的所有工件。为加速环境初始化,WeEnv即时启动环境并按需获取内容。在任务执行期间,WeEnv弹性供应CPU和内存,根据观测到的使用情况调整每个环境的配额,以适应变化的需求。与E2B、Docker和AgentENV相比,WeEnv将初始化时间减少了5.6-14.2倍,将其在迭代时间中的占比从最高53.4%降至9.1%。WeEnv已部署于微信的智能体强化学习。
英文摘要
Agentic reinforcement learning (RL) differs from conventional RL in that every task executes inside a complex environment, e.g., a virtual machine or a container. We find that agentic RL pays a heavy environment tax: a large share of the iteration time goes to the environment rather than to learning. The root cause is the lack of a full-lifecycle solution to environment management. We present WeEnv, which manages environments across packaging, initialization, and provisioning. WeEnv packages components as independently published layer groups and composes them at initialization, so that updating a component republishes one small group rather than every artifact containing it. To speed up environment initialization, WeEnv launches environments instantly and fetches contents on demand. During task execution, WeEnv provisions CPU and memory elastically, adjusting each environment's quota from its observed usage to fit the varying demands. WeEnv reduces the initialization by 5.6-14.2x over E2B, Docker, and AgentENV, cutting its share of the iteration time from up to 53.4% to 9.1%. WeEnv is deployed for agentic RL at WeChat.