用于LLM RL后训练中低成本故障复现与诊断的MoE代理模型
MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training
AI总结:
本文针对LLM RL后训练中故障复现成本高的问题,提出一种MoE代理模型构建方法,通过专家剪枝降低计算需求,可低成本复现故障并辅助诊断。
AI中文摘要:
大型语言模型(LLM)的强化学习(RL)后训练计算密集度高,且涉及复杂的系统流程,调试开销巨大。实际中,框架适配、数值精度、算子实现等因素会引发梯度溢出、损失发散等故障。直接在大型模型上复现此类故障需要大量时间和计算资源。本文系统分析了在华为昇腾(Huawei Ascend)平台上进行大规模RL训练时遇到的故障,总结了代表性故障类型,并确定了与故障复现相关的三个模型侧因素。基于这些因素,我们提出了一种用于低成本故障调查和辅助诊断的代理模型构建方法,该方法采用保留结构的、基于聚类的专家剪枝,以选择代表性专家,同时保留模型的主干架构、路由机制和基本任务能力。我们的实验结果表明,代理模型将加速器需求降低了50%-87.5%,单步NPU-小时成本最多降低了33.3倍,同时保留了主要的训练动态并复现了与原始模型一致的故障响应。总体而言,代理模型可作为RL后训练中故障复现、针对性验证和辅助诊断的低成本替代工具。
英文摘要:
Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive and involves complex system pipelines with substantial debugging overhead. In practice, factors such as framework adaptation, numerical precision, and operator implementation can cause failures, including gradient overflow and loss divergence. Reproducing such failures directly on large models requires considerable time and computational resources. This paper systematically analyzes failures encountered during large-scale RL training on the Huawei Ascend platform, summarizes representative failure types, and identifies three model-side factors relevant to fault reproduction. Based on these factors, we propose a proxy-model construction method for low-cost fault investigation and auxiliary diagnosis. It employs structure-preserving, clustering-based expert pruning to select representative experts while retaining the model's backbone architecture, routing mechanism, and basic task capabilities. Our experimental results show that the proxy models reduce accelerator requirements by 50%-87.5% and achieve up to a 33.3x reduction in per-step NPU-hour cost, while preserving major training dynamics and reproducing fault responses consistent with the original models. Overall, the proxy models can serve as low-cost surrogates for fault reproduction, targeted validation, and auxiliary diagnosis in RL post-training.