AI 中文总结
SubZero+是改进的SubZero框架,通过三种互补方式提升零阶优化稳定性,在1.3B至32B参数模型的SuperGLUE任务上,优于现有零阶基线,扩大稳定学习率范围,缩小与一阶方法差距且内存开销极小。
AI 中文摘要
零阶(ZO)优化可实现无需反向传播的大语言模型微调,但现有ZO方法的梯度估计器方差高,导致收敛不稳定且对学习率高度敏感。本文提出改进的SubZero框架SubZero+,通过三种互补方式提升稳定性:(i)在特定层的低秩子空间内进行多查询梯度估计,以降低方差且不出现多查询悖论;(ii)子空间Adam优化器,利用子空间内多查询梯度统计量执行自适应更新;(iii)基于QR的子空间构造的符号校正,确保哈尔分布投影矩阵,消除依赖实现的方向歧义。在SuperGLUE上对13亿至32亿参数的模型(含全参数微调与LoRA设置)开展实验,结果显示SubZero+始终优于现有ZO基线,扩大了稳定学习率范围,且以极小额外内存开销缩小了与一阶方法的差距。
英文摘要
Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization through a carefully designed dual low-dimensionality strategy: (i) multi-query forward-difference gradient estimation in periodically refreshed random subspaces to mitigate noise amplification in moment buffers, and (ii) Adam updates with periodic restarts performed directly in low-dimensional space rather than full-parameter space. In experiments, this dual design retains memory overhead comparable to momentum-free ZO methods while achieving stronger optimization performance than the evaluated ZO baselines. Theoretically, in the exact-directional limit, $K$-query averaging preserves conditional unbiasedness, while the coefficient estimator's covariance and mean-squared error, as well as query-induced second-moment inflation, scale exactly as $1/K$. Extensive experiments across SuperGLUE with models from 1.3B to 32B parameters under both full fine-tuning and LoRA schemes demonstrate consistent improvements over competing ZO methods. SubZero+ significantly narrows the performance gap with first-order optimization while preserving ZO's inference-time memory efficiency.