超越低秩参数化:通过梯度分解缩小LoRA与全微调之间的差距
Beyond Low-Rank Parameterization: Narrowing the Gap Between LoRA and Full Fine-Tuning via Gradient Decomposition
- Huazhong University of Science and Technology(华中科技大学)
- Hebei University of Technology(河北工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出GDLoRA,通过正交分解提取法向梯度更新基权重,在不增加内存下缩小LoRA与全微调的性能差距。
AI中文摘要:
低秩适配(LoRA)是一种广泛使用的参数高效微调(PEFT)方法,但与全微调(FFT)相比,仍可能存在性能差距。许多LoRA变体改进了低秩因子的初始化或优化。然而,在每个训练步骤中,它们的一阶权重空间方向受当前参数化的约束。我们刻画了相应的LoRA可访问梯度空间,并证明其与当前LoRA参数化诱导的切空间一致。这一刻画在当前模型参数下产生了全权重梯度的正交分解。我们将与该空间正交的分量称为法向梯度。基于此分解,我们提出了GDLoRA(梯度分解低秩适配)。GDLoRA从前向激活和反向信号中重建全权重梯度,提取其法向分量,并直接用该分量更新基权重,同时保留对LoRA因子的标准AdamW优化。在匹配适配器和优化器配置下,GDLoRA在不增加标准LoRA优化器状态内存预算的情况下,融入了互补的法向梯度。在自然语言理解、数学推理、常识推理和图像分类上的实验表明,GDLoRA持续优于LoRA,并缩小了与FFT的性能差距。代码可在该https URL获取。
英文摘要:
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full weight gradient at the current model parameters. We term the component orthogonal to this space the normal gradient. Based on this decomposition, we propose GDLoRA (Gradient-Decomposed Low-Rank Adaptation). GDLoRA reconstructs the full weight gradient from forward activations and backward signals, extracts its normal component, and directly updates the base weights with this component, while retaining standard AdamW optimization for the LoRA factors. GDLoRA incorporates complementary normal gradients without increasing standard LoRA's optimizer-state memory budget under matched adapter and optimizer configurations. Experiments on natural language understanding, mathematical reasoning, commonsense reasoning, and image classification show that GDLoRA consistently improves over LoRA and narrows the performance gap to FFT. The code is available at https://anonymous.4open.science/r/GDLoRA.