发表机构
Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出动量p维子空间信赖域方法(MpSub),用于大语言模型无导数微调,无需学习率,在匹配预算下达到与调优MeZO相当的准确率。
AI 中文摘要
大语言模型的全参数微调由于反向传播需要存储激活值和梯度,导致内存成本高昂。零阶优化通过从损失评估中估计更新方向来避免这一问题,但现有方法需要针对每个模型和任务调整敏感的学习率。我们提出了动量p维子空间信赖域方法(MpSub)。在每次迭代中,MpSub在p维子空间内进行搜索:一个方向保留最近接受步骤的历史动量,其余方向通过新的随机采样进行探索。子空间梯度通过中心差分估计,试验步长由线性信赖域模型计算,信赖域半径根据预测损失减少与实际损失减少的一致性自适应调整,从而消除了学习率。对于大语言模型微调,迭代内的评估共享一个小批量,方向从种子原地重新生成,仅使用前向传播。对于未正交化高斯方向下的光滑确定性目标,我们限定了有限差分误差,量化了子空间捕获的梯度能量,并证明了在保护半径更新下,$\lim_{k\to\infty} \\|\nabla f(x_k)\\|_2 = 0$几乎必然成立。在匹配的8,400次训练目标前向传播预算下,我们在CommitmentBank上微调了OPT-125M和OPT-350M。在两个模型规模下使用相同的预设参数,MpSub在三个随机种子上达到了0.673和0.690的平均测试准确率,与调优后的MeZO(0.685)相当,且无需任何学习率搜索。
英文摘要
Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation stores activations and gradients. Zeroth-order optimization avoids this by estimating update directions from loss evaluations, but existing methods require tuning a sensitive learning rate for each model and task. We propose the momentum $p$-dimensional subspace trust-region method (MpSub). At each iteration, MpSub searches within a $p$-dimensional subspace: one direction preserves historical momentum from the most recent accepted step, while the remaining directions explore via fresh random sampling. The subspace gradient is estimated by central differences, a trial step is computed from a linear trust-region model, and the trust-region radius adapts according to the agreement between predicted and observed loss reduction, eliminating the learning rate. For LLM fine-tuning, evaluations within an iteration share a minibatch, and directions are regenerated in place from seeds, using forward passes alone. For smooth deterministic objectives under unorthogonalized Gaussian directions, we bound the finite-difference error, quantify gradient energy captured by the subspace, and prove that $\lim_{k\to\infty} \|\nabla f(x_k)\|_2 = 0$ almost surely under a safeguarded radius update. Under a matched budget of 8,400 training-objective forward passes, we fine-tune OPT-125M and OPT-350M on CommitmentBank. With the same preset parameters at both model sizes, MpSub attains mean test accuracies of 0.673 and 0.690 over three seeds, matching tuned MeZO (0.685) without any learning-rate search.
Comments16 pages, 2 figures, 2 tables