LoRA-GA²:具有多步梯度自适应对齐的低秩适配
LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
浏览论文内容
中文总结 AI 辅助
本文提出LoRA-GA²算法,通过多步梯度的轻量探针、频谱感知秩分配及最优初始化,在保持LoRA效率的同时,在GLUE、GSM8K等基准上优于现有LoRA变体。
中文摘要 AI 辅助
低秩适配(LoRA)是一种针对大模型的突出微调方法,以更低的内存开销实现了具有竞争力的性能。然而,LoRA与全量微调之间仍存在持续的性能差距。近期研究尝试通过利用预训练权重的单步梯度近似,将LoRA更新与全量微调更新的主方向或固有维度对齐,以缩小该差距,但这些方法无法捕捉梯度的完整动态。本文提出LoRA-GA²,一种充分利用多步梯度信息的有效微调算法。具体而言,我们引入一种针对预训练权重多步梯度的轻量探针,该探针不产生额外GPU内存开销,仅带来可忽略的时间开销;我们还采用基于多步梯度的频谱感知、重要性驱动的秩分配及最优初始化。大量实验结果表明,LoRA-GA²在保持普通LoRA效率优势的同时,始终优于现有LoRA变体;例如,在GLUE基准上,LoRA-GA²较领先基线平均超出0.66个点,在GSM8K上超出最强基线1.03个点,在HumanEval上超出0.87个点。
英文摘要
Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent studies have sought to narrow this gap by employing one-step gradient approximations of pretrained weights to align LoRA updates with the principal directions or intrinsic dimensionalities of full fine-tuning updates. Nevertheless, these approaches fail to capture the full dynamics of the gradients. In this paper, we propose LoRA-GA$^2$, an effective fine-tuning algorithm that fully leverages multi-step gradient information. Specifically, we introduce a lightweight probe for multi-step gradients of pretrained weights that incurs no additional GPU memory cost and only marginal time overhead. We further employ a spectrum-aware, importance-based rank allocation and optimal initialization derived from multi-step gradients. Extensive experimental results demonstrate that LoRA-GA$^2$ consistently outperforms existing LoRA variants while preserving the efficiency advantages of vanilla LoRA. For instance, LoRA-GA$^2$ surpasses the leading baseline by an average of 0.66 points on the GLUE benchmark, and outperforms the strongest baseline by 1.03 points on GSM8K and 0.87 points on HumanEval, respectively.