arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全栈FP4:通过量化投影、优化器和注意力进行稳定的语言模型预训练

Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention

Siyu Ding, Mingchuan Ma, Jiabo Tong, Xingrun Xing, Ziming Wang, Guoqi Li

arXiv 2607.04422首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; Sichuan University; Beijing Zhongguancun Academy; Zhejiang University; Beijing Key Laboratory of Brain-Inspired General Intelligence Large Model; Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(中国科学院自动化研究所; 四川大学; 北京中关村科学院; 浙江大学; 北京脑科学与类脑研究中心通用人工智能大模型北京市重点实验室; 脑认知与脑机智能技术重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对4位预训练中优化器状态等未充分探索的问题,提出全栈FP4框架。通过模块精度策略解决稳定性瓶颈,如用LoRA-SVD抑制量化噪声,设计优化器变换,采用混合精度方案,实现稳定高效的端到端预训练。

AI 中文摘要

近期NVFP4预训练方法主要针对变压器线性层,优化器状态等在4位管道中未充分探索。这阻碍了稳定的全栈4位预训练,因三个核心模块有独特数值故障模式。我们提出全栈FP4,首个解决所有三个稳定性瓶颈的框架,通过模块精度策略,验证了可行的稳定端到端NVFP4语言模型预训练。

英文摘要

Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer computation, and low-precision attention forward--backward paths less explored. We present \textbf{Full-Stack FP4}, a modular NVFP4 framework with separate recipes for projections, AdamW states, Root/Muon computation, and attention. \textbf{LoRA-SVD} protects a compact projection subspace in BF16 while retaining full-shape NVFP4 computation, reducing the linear-only loss gap from \textbf{1.40\%} to \textbf{0.61\%}. An ordered square-root, tile-mean, and Hadamard pipeline enables stable NVFP4 AdamW momentum storage; shape-dependent coefficients and clipping stabilize direct NVFP4 Root iterations; and mixed-precision attention retains softmax-sensitive operations in BF16. On 3B pretraining with 64B tokens, BF16 and Full-Stack FP4 reach losses of \textbf{2.267} and \textbf{2.286}, a \textbf{0.838\%} gap. Their average zero-shot perplexities are 26.675 and 26.665, respectively, with Full-Stack FP4 averaging 0.10 percentage points lower in accuracy. Native four-block measurements on one RTX 5090 show 2.50--2.83$\times$ Root speedups over optimized BF16 and 37.9--42.5\% lower AdamW peak memory.

CommentsFix experiment bugs and some statement, update current developed hardware efficiency result on 5090

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑