MiMo-V2.6:将强化学习扩展至自我改进
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
浏览论文内容
中文总结 AI 辅助
本研究推出MiMo-V2.6系列,通过从三方面扩展RL计算,结合中间训练、分组智能体grading等技术,构建相关基础设施并开源部分内容,推动RL助力大型基础模型自我改进。
中文摘要 AI 辅助
强化学习(RL)是推动大型基础模型实现自我改进的核心训练范式。本报告介绍MiMo-V2.6系列,这是一个全模态家族,通过扩展RL计算推动模型智能的前沿发展。在RL训练前,我们对广泛的多模态语料库进行中间训练,以提供充足的探索空间,并在预训练的hybrid-SWA架构上构建坚实的基础设施,以支持后续的规模扩展。我们从三个维度扩展RL计算:(1)更大的批次和更高的吞吐量,采用异步训练,在上下文长度达1M时,每步消耗1568个样本和27-37亿个token;(2)更多样、更复杂的环境,涵盖代码、通用、视觉和网络领域,基于多种智能体 harness;(3)更多的 grader 计算,通过分组智能体 grading 为长视野任务生成更准确的奖励信号,并引导模型走向更短、更具token效率的解决方案。为在规模上保持训练稳定,我们冻结MoE路由器,并建立多层防御机制以抵御奖励黑客攻击。我们还构建了混合任务智能体RL的基础设施,包括统一的轨迹表示、高并发多框架rollout、解耦的控制与数据平面,以及训练-推理一致性。我们开源了训练动态、RL环境和RL框架,以促进可复现性和对规模化RL及模型自我改进的进一步研究。
英文摘要
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.
发表机构
- Xiaomi(小米)
机构由 AI 辅助整理,请以论文原文为准。