arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08049cs.LG

低秩适应的黎曼几何

A Riemannian Geometry for Low-rank Adaptation

Shoichiro Takeda, Shin'ya Yamaguchi, Satoshi Suzuki, Yasunori Akagi

首次发表
浏览论文内容

中文总结 AI 辅助

针对低秩适应(LoRA)提出一种新的黎曼度量,使权重更新更接近全量微调的方向和结果,理论证明并实验验证其高效性。

中文摘要 AI 辅助

低秩适应(LoRA)作为一种参数高效的微调技术被广泛用于预训练深度神经网络,它通过一个低秩矩阵 $BA^\top$ 来近似全量微调中的权重更新。这种参数化导致了等价关系 $(B, A) \sim (BG^{-1}, AG^\top)$,其中 $G$ 为任意可逆矩阵,因为 $BA^\top = BG^{-1}(AG^\top)^\top$,因此两对参数产生相同的损失值。这一关系诱导出一个商流形,其中所有 $G$ 对应的矩阵对 $(BG^{-1}, AG^\top)$ 被视为等同,从而消除了损失值不变的多余方向。为尊重该流形的几何结构,原始搜索空间被赋予一个在该等价关系下不变的黎曼度量。这样的度量在每一步梯度更新时引入预条件,并确保每次通过 LoRA 进行的权重更新都会改变损失值,从而实现高效优化。本文提出一种专门针对 LoRA 设计的新黎曼度量,以在权重层面缩小与全量微调的差距。我们从理论上证明,采用该度量诱导的预条件的 LoRA 在每次迭代中满足以下两个性质:(i)权重更新在 LoRA 参数化所允许的一阶权重变化子空间内,沿最接近全量微调梯度的方向进行;(ii)更新后的权重矩阵在 Frobenius 范数意义上比采用传统预条件或无预条件的 LoRA 更新后的权重矩阵更接近全量微调的权重矩阵。这些理论见解表明,我们的预条件使 LoRA 更好地逼近全量微调,从而实现更高效的优化。实验在语言和视觉领域的微调任务上验证了我们的预条件对 LoRA 的有效性和高效性。

英文摘要

Low-rank adaptation (LoRA) is widely used as a parameter-efficient fine-tuning technique for pre-trained deep neural networks, which approximates the weight update via full fine-tuning by a low-rank matrix $BA^\top$. This parameterization leads to the equivalence relation $(B, A) \sim (BG^{-1}, AG^\top)$ for any invertible matrix $G$ because $BA^\top = BG^{-1}(AG^\top)^\top$ and thus both pairs yield the same loss value. This relation induces a quotient manifold where matrices $(BG^{-1}, AG^\top)$ for all $G$ are identified, eliminating redundant directions along which the loss value remains unchanged. To respect the geometry of this manifold, the original search space is endowed with a Riemannian metric that is invariant under the equivalence relation. Such a metric induces preconditioning at each gradient step and ensures that each weight update via LoRA changes the loss value, leading to efficient optimization. In this paper, we propose a new Riemannian metric that is specifically tailored to LoRA to close the gap to full fine-tuning at the weight level. We theoretically show that LoRA with our preconditioning induced by this metric satisfies the following two properties at each iteration: (i) The weight update follows the direction closest to the gradient of full fine-tuning within the subspace of first-order weight changes allowed by the LoRA parameterization. (ii) The updated weight matrix is closer in Frobenius norm to that of full fine-tuning than the updated weight matrices of LoRA with conventional preconditioning and without preconditioning. These theoretical insights suggest that our preconditioning makes LoRA better approximate full fine-tuning, thereby leading to more efficient optimization. Experiments show the effectiveness and efficiency of our preconditioning for LoRA on fine-tuning tasks with language and vision domains.

发表机构

  • NTT, Inc.(日本电报电话公司)

机构由 AI 辅助整理,请以论文原文为准。

↑