arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VenusRL:具有优先级调度和可扩展交互的完全解耦智能体强化学习系统

VenusRL: A Fully Disaggregated Agentic RL System with Priority Scheduling and Scalable Interaction

Mingjun Zhang, Yucheng Li, Menghao Zhang, Shuyong Zhu, Ping Zhang, Xiaohe Hu, Jun Chen, Zhixin Wang, Xutong Wang, He Liu, Yanmin Jia, Shengrong Zhu, Peng Sun, Mingjie Zhang, Liming Liu, Jinlong Hou, Yuan Cheng, Yujun Zhang

arXiv 2610.03286首次发表:更新:

发表机构

Institute of Computing Technology, Chinese Academy of Sciences; Beihang University; Infrawaves; Shanghai Innovation Institute; Shanghai Qiji Zhifeng Co., Ltd.(中国科学院计算技术研究所; 北京航空航天大学; Infrawaves; 上海创新研究院; 上海奇绩智峰有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体强化学习训练中吞吐量受限于慢轨迹和沙箱内存浪费两大瓶颈,提出完全解耦系统VenusRL,通过优先级调度和内存共享实现4.24倍加速及89%成本降低。

AI 中文摘要

智能体强化学习(Agentic Reinforcement Learning, RL)通过智能体与外部工具环境的多轮交互来训练大型语言模型(LLM)智能体。其多轮特性暴露了现有智能体RL框架未解决的两个系统级瓶颈。首先,端到端训练吞吐量受限于完成最慢的轨迹,然而仅优化每个GPU的利用率会使回滚进度分散到许多组中,延迟足够多组的完成以解锁下一步训练。其次,工具沙箱根据其声明的内存上限静态过度配置,导致大部分物理内存闲置,同时从同一提示启动的沙箱复制几乎相同的状态。我们提出了VenusRL,一个完全解耦的智能体RL系统,解决了这两个瓶颈。VenusRL的优先级感知动作调度器使用长度预测启发式方法识别最有可能解锁下一步训练的样本组,并在批处理准入、KV缓存驻留和跨工作节点请求编排中将其置于其他组之前。VenusRL的环境资源管理器结合了内存感知准入阈值和模板键控页面共享池,在保持严格内存隔离的同时,通过写保护页表项别名和写时复制,在每节点打包更多沙箱。在代表性智能体RL工作负载上,VenusRL相比最先进的基线实现了高达4.24倍的端到端训练加速,并将环境成本降低了高达89%。

英文摘要

Agentic Reinforcement Learning (RL) trains LLM agents through multi-turn interactions with external tool environments. Its multi-turn nature exposes two system-level bottlenecks unaddressed by existing agentic RL frameworks. First, end-to-end training throughput is constrained by the slowest trajectories to complete, yet optimizing per-GPU utilization alone scatters rollout progress across many groups, delaying the completion of enough groups to unblock the next training step. Second, tool sandboxes are statically over-provisioned by their declared memory ceilings, leaving most physical memory stranded while replicating near-identical state across sandboxes launched from the same prompt. We present VenusRL, a fully disaggregated agentic RL system that addresses both bottlenecks. VenusRL's priority-aware action scheduler uses length-prediction heuristics to identify sample groups whose completion is most likely to unblock the next training step, and pushes them ahead of others across batch admission, KV Cache residency, and cross-worker request orchestration. VenusRL's environment resource manager combines a memory-aware admission threshold with a template-keyed page-sharing pool, packing more sandboxes per node while preserving strict memory isolation via write-protected page table entry aliasing and copy-on-write. Across representative agentic RL workloads, VenusRL achieves up to 4.24x end-to-end training speedup over state-of-the-art baselines and reduces environment cost by up to 89%.

Comments18 pages, 19 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑