发表机构
University of California, Berkeley; Nubank(加州大学伯克利分校; 努班克)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对AdamW优化器状态量化的误差问题,提出ZIP-SR和ZE-EDEN两种4位量化方案,在1.3亿至27亿参数的预训练中缩小了与32位AdamW的损失差距,微调时性能优于TorchAO。
AI 中文摘要
对AdamW的优化器状态进行量化可减少持久存储,但量化误差会通过矩递推传播并干扰后续的自适应更新。我们从\textit{舍入空间}(量化器在相邻重构电平间选择的坐标)的角度重新设计AdamW的4位优化器状态量化。针对二阶矩,对零值相邻量化单元的局部分析表明,小的平均状态误差未必意味着下一步小的平均预条件误差。一维二次构造进一步显示,在状态空间和预条件空间舍入下,优化动态存在质的差异。这些结果催生了零包含型预条件空间随机舍入(\textbf{ZIP-SR}),该方法在二阶矩码本中保留零值,并在预条件空间中计算随机舍入概率。作为补充方案,零排除型EDEN校准(\textbf{ZE-EDEN})采用排除零值的二阶矩码本,并对量化后的二阶矩块进行重缩放,以缓解正量化下限导致的预条件失真。两种配置均对一阶矩使用4位NormalFloat(NF4),并在训练的最后10%阶段对LM头一阶矩进行针对性随机舍入。在从\textbf{1.3亿}到\textbf{27亿}参数的GPT和Llama风格预训练实验中,两种方法在所有评估的模型规模下均缩小了TorchAO 4位AdamW与32位AdamW的平均验证损失差距,报告的最大差距缩小幅度达\textbf{70%}。在全参数监督微调中,两种方案均实现了比TorchAO更低的验证损失,同时在下游任务上接近32位AdamW的性能。
英文摘要
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error need not imply small mean preconditioner error at the next step. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (\textbf{ZIP-SR}), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (\textbf{ZE-EDEN}) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10\% of training. Across GPT- and Llama-style pretraining experiments ranging from \textbf{130M} to \textbf{2.7B} parameters, both methods reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size, with the largest reported gap reduction reaching \textbf{70\%}. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks.
Comments23 pages