arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

梯度与自然梯度之间:LoRA初始化的连续统

Between Gradient and Natural Gradient: A Continuum of LoRA Initializations

Dianze Liu, Farshid Ghezelbash

arXiv 2607.26247首次发表:更新:

AI 中文总结

该研究提出统一LoRA(ULoRA)框架,将现有LoRA初始化方法整合为连续统,发现最佳预条件强度依赖任务,其变体ULoRA-Auto无需额外搜索即可接近上界性能,提升了LoRA的适配性。

AI 中文摘要

低秩适配(LoRA)以仅为全参数微调一小部分的成本微调大型预训练模型,但其性能强烈依赖于适配器的初始化方式。近期方案从下游损失梯度初始化适配器:一些将原始梯度投影到其主方向,另一些则先通过损失曲率估计对其进行白化处理。我们表明,这些看似不同的方法属于单一连续统:一个由谱白化指数和类Adam对角指数控制的双参数预条件梯度初始化族,我们称之为统一LoRA(ULoRA)。在完整学习率搜索下遍历该族,我们发现没有单一固定预条件强度占优:最佳工作点依赖于任务,且常严格位于族内,远离已发表的端点。将调优后的ULoRA配置视为该族的上界,在RoBERTa-base上的全部5个GLUE任务上匹配或超越全微调性能,在LLaMA-2-7B上的GSM8K任务上与最强基准相当。我们可部署的无搜索变体ULoRA-Auto从测量的谱统计中选择每层指数,无需额外搜索成本即可接近该上界,在可部署LoRA方法中排名靠前或接近靠前。我们的结果表明,LoRA初始化和曲率预条件的原则性设计空间应被视为可调维度,而非固定设计决策。

英文摘要

Low-rank adaptation (LoRA) fine-tunes large pretrained models at a fraction of the cost of full fine-tuning, but its performance depends strongly on how the adapters are initialized. Recent schemes initialize the adapters from the downstream loss gradient: some project the raw gradient onto its top directions, while others first whiten it with an estimate of the loss curvature. We show that these seemingly distinct methods are points on a single continuum: a two-parameter family of preconditioned gradient initializations, which we call Unified LoRA (ULoRA), governed by a spectral whitening exponent and an Adam-like diagonal exponent. Sweeping this family under a full learning-rate search, we find that no single fixed preconditioning strength dominates: the best operating point is task-dependent and frequently lies strictly inside the family, away from the published endpoints. Treated as an upper bound of this family, a tuned ULoRA configuration matches or exceeds full fine-tuning on all five GLUE tasks with RoBERTa-base and is competitive with the strongest baselines on GSM8K with LLaMA-2-7B. Our deployable, search-free variant, ULoRA-Auto, selects per-layer exponents from measured spectral statistics, approaches this upper bound at no additional search cost, and ranks at or near the top among deployable LoRA methods. Our results show that a principled design space for LoRA initialization and curvature preconditioning should be treated as a tunable dimension rather than a fixed design decision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑