AI 中文总结
研究针对万亿参数规模MoE模型全参数训练后优化的系统挑战,在升腾NPU超级集群上以DeepSeek-V4为目标工作负载开发分层优化框架,建立CPT和SFT工作流程,提升模型性能,推动复杂推理前沿模型系统发展。
AI 中文摘要
万亿参数规模的混合专家(MoE)模型的全参数训练后优化给大规模分布式训练带来了巨大的系统级挑战,包括严重的内存压力、通信开销和内核执行效率低下等问题。虽然大多数大规模语言模型(LLM)训练系统基于GPU集群构建,但本报告展示了在升腾NPU超级集群上的端到端优化实践。以DeepSeek-V4模型系列为目标工作负载,开发了一个涵盖模型级并行、计算-通信编排和底层内核执行的分层优化框架。优化后的系统实现了34.22%的模型浮点运算利用率(MFU),比开源基线方法提高了2.93倍,同时保持了训练稳定性。在此基础上,为复杂运筹学(OR)任务建立了CPT和SFT工作流程。使用DeepSeek-V4-Flash开发了面向OR的CPT和SFT数据管道,将收集的领域资源与求解器验证的合成优化文档相结合。得到的数据集包含10K高质量SFT样本,跨越四个任务类别和三种问题表示形式。该专门模型在评估模型中达到了最高的平均零样本Pass@1分数,达到71.81%,分别比GPT-5.4-Mini和基础DeepSeek-V4-Flash模型高出3.98和11.27个百分点。总体而言,这项工作展示了一条从升腾基础设施上高效的万亿参数模型训练后优化到用于求解器基础数学建模的领域专用Flash模型的全栈路径,推动了复杂推理的前沿模型系统发展。
英文摘要
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
Comments73 pages, 22 figures, 20 tables