arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.09034cs.LGcs.AI

FZOO:面向Adam级速度的大语言模型微调快速零阶优化器

FZOO: Fast Zeroth-Order Optimizer for Fine-Tuning Large Language Models towards Adam-Scale Speed

  • Xi’an Jiaotong University(西安交通大学)
  • SGIT AI Lab(SGIT AI实验室)
  • A*STAR(新加坡科技研究局)

机构由 AI 辅助整理,请以论文原文为准。

Sizhe Dang, Yangyang Guo, Yanjun Zhao, Haishan Ye, Xiaodong Zheng, Guang Dai, Ivor Tsang

更新

AI总结:

针对大语言模型微调中的内存瓶颈,提出快速零阶优化器FZOO,通过批量单侧估计和Rademacher扰动减少前向传播次数并加速计算,在准确率和收敛速度上显著优于MeZO,实现单GPU全参数微调。

AI中文摘要:

微调大语言模型(LLMs)常面临GPU内存瓶颈:Adam等一阶优化器的反向传播使内存使用量增至推理水平的10倍以上(例如OPT-30B需633 GB)。零阶(ZO)优化器仅通过前向传播估计梯度以避免此开销,但MeZO等现有方法通常需要更多步数才能收敛。ZO中速度与内存的权衡能否得到根本性改善?Normalized-SGD展现出强劲的实证性能,且比Adam具有更高的内存效率。基于此,我们提出FZOO,一种面向Adam级速度的快速零阶优化器。FZOO通过采用批量单侧估计(基于批量损失的标准差调整步长),减少了收敛所需的总前向传播次数。它还利用Rademacher随机向量扰动结合CUDA的并行处理,加速了每批次计算。在RoBERTa-large、OPT(350M-66B)、Phi-2和Llama3等多种模型及11项任务上的广泛实验验证了FZOO的有效性。平均而言,FZOO在准确率上比MeZO高出3%,同时所需的前向传播次数减少3倍。对于RoBERTa-large,FZOO实现了平均5.6%的准确率提升,且前向传播次数比MeZO减少18倍,收敛速度与Adam相当。我们还提供了理论分析,证明FZOO与Normalized-SGD更新规则的正式等价性及其收敛保证。FZOO可平滑集成至PEFT技术中,实现更大的内存节省。总体而言,我们的结果使单GPU、高速、全参数微调变得切实可行,并为内存高效的预训练指明了未来工作方向。

英文摘要:

Fine-tuning large language models (LLMs) often faces GPU memory bottlenecks: the backward pass of first-order optimizers like Adam increases memory usage to more than 10 times the inference level (e.g., 633 GB for OPT-30B). Zeroth-order (ZO) optimizers avoid this cost by estimating gradients only from forward passes, yet existing methods like MeZO usually require many more steps to converge. Can this trade-off between speed and memory in ZO be fundamentally improved? Normalized-SGD demonstrates strong empirical performance with greater memory efficiency than Adam. In light of this, we introduce FZOO, a Fast Zeroth-Order Optimizer toward Adam-Scale Speed. FZOO reduces the total forward passes needed for convergence by employing batched one-sided estimates that adapt step sizes based on the standard deviation of batch losses. It also accelerates per-batch computation through the use of Rademacher random vector perturbations coupled with CUDA's parallel processing. Extensive experiments on diverse models, including RoBERTa-large, OPT (350M-66B), Phi-2, and Llama3, across 11 tasks validate FZOO's effectiveness. On average, FZOO outperforms MeZO by 3 percent in accuracy while requiring 3 times fewer forward passes. For RoBERTa-large, FZOO achieves average improvements of 5.6 percent in accuracy and an 18 times reduction in forward passes compared to MeZO, achieving convergence speeds comparable to Adam. We also provide theoretical analysis proving FZOO's formal equivalence to a normalized-SGD update rule and its convergence guarantees. FZOO integrates smoothly into PEFT techniques, enabling even larger memory savings. Overall, our results make single-GPU, high-speed, full-parameter fine-tuning practical and point toward future work on memory-efficient pre-training.

↑