arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2312.15184cs.LG

ZO-AdaMU优化器:在零阶优化中通过动量和不确定性自适应扰动

ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-order Optimization

  • Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
  • Peng Cheng Laboratory(鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

Shuoran Jiang, Qingcai Chen, Youchen Pan, Yang Xiang, Yukang Lin, Xiangping Wu, Chuanyi Liu, Xiaobao Song

更新

AI总结:

针对零阶优化中MeZO存在的振荡、过拟合和收敛慢问题,提出ZO-AdaMU,通过在模拟扰动上应用动量自适应,提升收敛稳定性与速度,并在LLM微调中取得更好泛化。

AI中文摘要:

降低大模型全参数训练中的内存需求已成为一个热门研究领域。MeZO通过仅使用前向传播的零阶SGD优化器(ZO-SGD)对大型语言模型(LLMs)进行微调,展示了在GPU内存使用与推理相同的情况下具有出色的性能。然而,MeZO中用于梯度估计的模拟扰动随机近似会导致严重的振荡,并产生大量的时间开销。此外,在没有动量正则化的情况下,MeZO表现出严重的过拟合问题。最后,ZO-SGD上与扰动无关的动量并未提高收敛速度。本研究提出了ZO-AdaMU,通过在其随机近似中利用动量自适应模拟扰动来解决上述问题。与现有的自适应动量方法不同,我们将动量重新定位在随机梯度近似中的模拟扰动上。我们的收敛性分析和实验证明,这是提高ZO-SGD收敛稳定性和收敛速度的更好方法。大量实验表明,与MeZO及其动量变体相比,ZO-AdaMU在各种NLP任务的LLMs微调中具有更好的泛化性能。

英文摘要:

Lowering the memory requirement in full-parameter training on large models has become a hot research area. MeZO fine-tunes the large language models (LLMs) by just forward passes in a zeroth-order SGD optimizer (ZO-SGD), demonstrating excellent performance with the same GPU memory usage as inference. However, the simulated perturbation stochastic approximation for gradient estimate in MeZO leads to severe oscillations and incurs a substantial time overhead. Moreover, without momentum regularization, MeZO shows severe over-fitting problems. Lastly, the perturbation-irrelevant momentum on ZO-SGD does not improve the convergence rate. This study proposes ZO-AdaMU to resolve the above problems by adapting the simulated perturbation with momentum in its stochastic approximation. Unlike existing adaptive momentum methods, we relocate momentum on simulated perturbation in stochastic gradient approximation. Our convergence analysis and experiments prove this is a better way to improve convergence stability and rate in ZO-SGD. Extensive experiments demonstrate that ZO-AdaMU yields better generalization for LLMs fine-tuning across various NLP tasks than MeZO and its momentum variants.

↑