arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于LoRA微调的Stiefel流形上无回缩优化

Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning

Yuan Zhang, Jiang Hu, Zhijian Lai, Lin Lin, Zaiwen Wen

arXiv 2607.25299首次发表:更新:

发表机构

Center for Data Science, Peking University; Yau Mathematical Sciences Center, Tsinghua University; Beijing International Center for Mathematical Research, Peking University; Department of Mathematics, University of California, Berkeley; Center for Machine Learning Research and Changsha Institute for Computing and Digital Economy, Peking University(北京大学数据科学中心; 清华大学丘成桐数学科学中心; 北京大学北京国际数学研究中心; 美国加州大学伯克利分校数学系; 北京大学机器学习研究中心和长沙计算与数字经济研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对Stiefel流形优化中现有方法的问题,提出无回缩且无惩罚参数算法,利用相关特性建立收敛保证。将LoRA微调问题转为流形优化问题,引入Manifold-LoRA,经实验验证其在加速训练及下游性能方面的优势。

AI 中文摘要

Stiefel流形上的优化在各种机器学习任务中起着重要作用。现有方法要么使用回缩算子,对大规模矩阵需要昂贵的正交归一化,要么采用依赖仔细步长选择和惩罚参数调整的着陆方法。为应对这些挑战,我们提出一种直接着陆于流形的无回缩且无惩罚参数的算法。通过利用二次惩罚函数的强凸性和Stiefel流形的近端光滑性,在恒定和递减步长下建立了具有最佳已知迭代复杂度的全局收敛保证。然后,将大语言模型的低秩适应(LoRA)微调问题重新表述为流形优化问题,引入几何加速适应的Manifold-LoRA。该方法采用提出的着陆技术和精心设计的步长策略来加速训练过程。在基准数据集上的数值实验证明了该方法的效率和强大的下游性能。

英文摘要

Optimization over the Stiefel manifold plays a significant role in various machine learning tasks. Existing methods either use the retraction operators, requiring costly orthonormalization for large-scale matrices, or employ landing methods that rely on careful step size selection and penalty parameter tuning. To address these challenges, we propose a retraction-free and penalty parameter-free algorithm that directly lands on the manifold. By leveraging the strongly-convex-like property of the quadratic penalty function and the proximal smoothness of the Stiefel manifold, we establish global convergence guarantees with the best-known iteration complexities under both constant and diminishing step sizes. Then, we reformulate the low-rank adaptation (LoRA) fine-tuning problem for large language models as a manifold optimization problem, introducing Manifold-LoRA for geometry-accelerated adaptation. This approach employs the proposed landing technique and a carefully designed step size strategy to accelerate the training process. Numerical experiments on benchmark datasets demonstrate the efficiency and strong downstream performance of the proposed method.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑