arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31036cs.LG

归一化低秩适配

Normalized Low-Rank Adaptation

Jiale Kang, Ziyin Yue, Zheng Zhan, Yangyi Huang, Weiyang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LoRA训练动态正则化不足的问题,提出归一化低秩适配(NoRA),通过归一化下投影矩阵提升LoRA性能,在多任务中加速收敛、改善稳定性并缓解灾难性遗忘,且无额外计算开销。

中文摘要 AI 辅助

尽管低秩适配(LoRA)被广泛用于参数高效的模型适配,但其训练动态的正则化以实现稳定且有效的优化仍未得到充分探索。由于LoRA将上投影初始化为零,其早期优化动态主要由下投影决定。基于这一观察,我们提出了归一化低秩适配(NoRA),这是一种简单却有效的方法,在训练过程中对下投影矩阵进行归一化。我们进一步表明,这种归一化仅在初始化时应用即可,无需在整个训练过程中重复归一化就能改进标准LoRA。在预训练、监督微调以及强化学习任务中,NoRA始终能加快收敛速度、提升性能与训练稳定性,并缓解灾难性遗忘。这些优势无需额外的可训练参数,也不增加推理时的计算量,使NoRA成为LoRA的一种简单且适用范围广泛的增强方案。

英文摘要

While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. Building on this observation, we introduce Normalized Low-Rank Adaptation (NoRA), a simple yet effective method that normalizes the down-projection matrices during training. We further show that the same normalization can be applied only at initialization, improving standard LoRA without requiring repeated normalization throughout training. Across pretraining, supervised finetuning, and reinforcement learning, NoRA consistently accelerates convergence, improves performance and training stability, and mitigates catastrophic forgetting. These benefits require neither additional trainable parameters nor inference-time computation, making NoRA a simple and broadly applicable enhancement to LoRA.

发表机构

  • Yuanshi Intelligence(元氏智能)
  • The Chinese University of Hong Kong(香港中文大学)
  • Microsoft Research(微软研究院)
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

↑