发表机构
Institute of Automation, Chinese Academy of Sciences; Sichuan University; Beijing Zhongguancun Academy; Zhejiang University; Beijing Key Laboratory of Brain-Inspired General Intelligence Large Model; Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(中国科学院自动化研究所; 四川大学; 北京中关村科学院; 浙江大学; 北京脑科学与类脑研究中心通用人工智能大模型北京市重点实验室; 脑认知与脑机智能技术重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对4位预训练中优化器状态等未充分探索的问题,提出全栈FP4框架。通过模块精度策略解决稳定性瓶颈,如用LoRA-SVD抑制量化噪声,设计优化器变换,采用混合精度方案,实现稳定高效的端到端预训练。
AI 中文摘要
近期NVFP4预训练方法主要针对变压器线性层,优化器状态等在4位管道中未充分探索。这阻碍了稳定的全栈4位预训练,因三个核心模块有独特数值故障模式。我们提出全栈FP4,首个解决所有三个稳定性瓶颈的框架,通过模块精度策略,验证了可行的稳定端到端NVFP4语言模型预训练。
英文摘要
Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer computation, and low-precision attention forward--backward paths less explored. We present \textbf{Full-Stack FP4}, a modular NVFP4 framework with separate recipes for projections, AdamW states, Root/Muon computation, and attention. \textbf{LoRA-SVD} protects a compact projection subspace in BF16 while retaining full-shape NVFP4 computation, reducing the linear-only loss gap from \textbf{1.40\%} to \textbf{0.61\%}. An ordered square-root, tile-mean, and Hadamard pipeline enables stable NVFP4 AdamW momentum storage; shape-dependent coefficients and clipping stabilize direct NVFP4 Root iterations; and mixed-precision attention retains softmax-sensitive operations in BF16. On 3B pretraining with 64B tokens, BF16 and Full-Stack FP4 reach losses of \textbf{2.267} and \textbf{2.286}, a \textbf{0.838\%} gap. Their average zero-shot perplexities are 26.675 and 26.665, respectively, with Full-Stack FP4 averaging 0.10 percentage points lower in accuracy. Native four-block measurements on one RTX 5090 show 2.50--2.83$\times$ Root speedups over optimized BF16 and 37.9--42.5\% lower AdamW peak memory.
CommentsFix experiment bugs and some statement, update current developed hardware efficiency result on 5090