发表机构
Li Auto Inc(理想汽车公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
马赫思维-4-闪是含3B激活参数的35B参数专家混合智能体模型,通过训练后优化及可扩展交互环境提升性能。其流程含三个阶段,在多个基准测试中成绩优异,以低推理成本领先或匹配数倍激活规模模型。
AI 中文摘要
我们提出了马赫思维-4-闪,这是一个具有3B激活参数的35B参数的专家混合(MoE)智能体模型。仅通过训练后优化而不扩大预训练计算规模,该模型就达到了与100B参数级模型相当或超越其性能。通过为大规模强化学习引入可扩展的智能体交互环境,该模型在实际应用任务中取得了显著的性能提升。我们的流程包括三个阶段:统一的强化学习/优化训练基础设施,多领域强化学习专家并行训练后融合,以及混合中位数长度策略优化。马赫思维-4-闪在多个基准测试中取得了优异成绩。
英文摘要
We present Mach-Mind-4-Flash, a 35B-parameter Mixture-of-Experts (MoE) agentic model with 3B activated parameters. Through post-training optimization alone without scaling pre-training compute, the model achieves performance on par with or surpassing that of 100B-parameter-class models. By introducing scalable agentic interaction environments for large-scale reinforcement learning, the model attains significant performance gains on real-world application tasks. Our pipeline comprises three stages: (1) a unified RL/OPD training infrastructure with dynamic multi-teacher scheduling and operator-level acceleration, delivering 17\% end-to-end training speedup; (2) multiple domain-specific RL experts trained in parallel across Reasoning, General, and Agent tracks, then fused into a single generalist via Multi-Teacher On-Policy Distillation (MOPD) -- a routed reverse-KL objective that eliminates the see-saw degradation of mixed-reward RL; (3) Hybrid Median-length Policy Optimization (HMPO), a single-stage token-efficiency method that compresses reasoning chains by 19--46\% with $\le$0.7 percentage-point accuracy loss. Mach-Mind-4-Flash scores 92.70 on AIME'26, 82.82 on IFBench, 80.74 on Behavioral-SafetyBench, 75.80 on BFCL-v4, 72.31 on BrowseComp-zh, and 84.20 on ClawBench -- leading or matching models with 10--30$\times$ its activated size at a fraction of the inference cost.