发表机构
RadixArk(睿数方舟)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Miles v0.1是一个生产级全栈后训练系统,通过可验证、干净且可定制的组件设计,支持大规模RL训练,并已在GLM-5.2模型上验证了高效性。
AI 中文摘要
我们推出了Miles v0.1,这是一个用于前沿后训练的全栈、生产就绪系统。基于slime的简洁设计,Miles将强化学习(RL)训练循环的每个阶段都围绕一个原则进行设计:组件应经过验证、干净且可定制。以准确性、效率、可靠性和可扩展性为首要目标,Miles旨在让前沿规模的RL对研究人员和企业都易于使用。本报告端到端地介绍了该系统:基于SGLang的rollout引擎、可选择两种后端(NVIDIA Megatron-LM和PyTorch FSDP)的训练器,以及针对不同部署拓扑的三种权重同步传输方式。除了全参数RL,Miles还支持LoRA RL、在策略蒸馏、监督微调以及真正的在策略rollout-训练对齐,并将相同的架构扩展到扩散模型。最后,我们以一个端到端案例研究作结:在GLM-5.2 744B-A40B模型上,针对终端使用编码任务进行完全异步的智能体RL,运行在64块NVIDIA GB300 GPU上,前30个测量步骤的中位步时长为263秒。Miles已在https URL开源,项目网站位于https URL。
英文摘要
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.
Comments34 pages, 5 figures, 9 tables. Technical report