arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TaRA:训练感知型低秩适配初始化

TaRA: Training-Aware Low-Rank Adaptation Initialization

Taehyeon Kim, Eunhyeok Park

arXiv 2609.02639首次发表:更新:

发表机构

Pohang University of Science and Technology (POSTECH)(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LoRA初始化受低秩分解信息瓶颈影响性能敏感的问题,提出TaRA方法,通过初始化LoRA使低秩因子梯度近似全秩权重梯度,在可忽略计算开销下提升梯度保真度,且性能优于现有SOTA方法。

AI 中文摘要

低秩适配(LoRA)已成为参数高效微调(PEFT)的事实上的标准,但其性能对初始化高度敏感,这是由低秩分解带来的信息瓶颈导致的。现有方法试图通过利用预训练权重、激活函数或梯度的主成分来构建高质量的LoRA初始化,但这些方法并未直接考虑全秩模型的训练动态。在本文中,我们提出训练感知型低秩适配初始化(TaRA),该方法对LoRA进行初始化,使得低秩因子产生的梯度能紧密近似对应全秩权重矩阵的梯度。TaRA由数学公式推导而来,在训练开始时提高了梯度保真度,同时引入的计算开销可忽略不计。在多样且具有挑战性的微调任务中,TaRA始终优于现有的最先进方法,为有效的LoRA初始化提供了一种简单、鲁棒且可扩展的解决方案。

英文摘要

Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches attempt to construct high-quality LoRA initializations by exploiting principal components of pretrained weights, activations, or gradients. However, these methods do not directly account for the training dynamics of the full-rank model. In this paper, we propose Training-aware Low-Rank Adaptation Initialization (TaRA), a method that initializes LoRA such that the gradients induced by the low-rank factors closely approximate the gradient of the corresponding full-rank weight matrix. Derived from a mathematical formulation, TaRA improves gradient fidelity at the start of training while introducing negligible computational overhead. Across diverse and challenging fine-tuning tasks, TaRA consistently outperforms prior state-of-the-art methods, establishing a simple, robust, and scalable solution for effective LoRA initialization.

CommentsAccepted to the EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑