arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Molt:用于智能体强化学习的可扩展原生 PyTorch 训练框架

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Jian Hu, Huiying Li, Hao Zhang, Binfeng Xu, Yifan Zhang, Shaokun Zhang, Hemil Desai, Michael Demoret, Pavlo Molchanov, Jan Kautz, Yi Dong

arXiv 2607.21653首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对智能体强化学习中算法修改成本高的问题,提出原生 PyTorch 训练框架 Molt,其代码简洁,智能体是普通程序,通过异步循环训练策略,在全异步协议下性能与先进堆栈相当,且开源提供资源。

AI 中文摘要

智能体强化学习研究不断有算法修改、新估计器、新管道阶段和新的展开方案。在主流框架中,每次改变都要经过训练器、分布式后端和展开胶水等多层,成本很高。Molt 是一个原生 PyTorch 训练框架,旨在降低成本。其代码库紧凑且干净,便于研究人员理解和 AI 编码助手处理。智能体是普通程序,通过异步循环训练多模态和专家混合策略,且从不训练未生成的令牌。在匹配的全异步协议下,Molt 与基于 Megatron 的先进堆栈在统计上相当。Molt 开源并提供相关资源。

英文摘要

Agentic reinforcement learning requires rapid experimentation with agents and learning algorithms, yet large policies and long, multimodal trajectories demand substantial distributed infrastructure. We present MOLT, a lightweight, PyTorch- and Hugging Face-native framework that brings these goals together through four contributions. MOLT combines direct loading of Hugging Face models with experimentally validated trillion-parameter scalability in approximately 9.2K lines of framework code. Unified OpenAI- and Anthropic-compatible interfaces integrate existing agents with automatic handling of context compaction. Fully asynchronous training overlaps agent rollouts and policy optimization, accommodating variable agent execution times. Distributed experience storage removes centralized rollout-memory bottlenecks for long, multimodal trajectories. We experimentally validate the complete RL training pipeline on a one-trillion-parameter policy and demonstrate sustained learning with a 30B mixture-of-experts agent, establishing MOLT as a lightweight foundation for large-scale agentic RL research.

Commentsupdate tech report

DOI:10.13140/RG.2.2.23375.65447

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑