arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LoRA 微调格局的细粒度分析及其对数据选择的影响

A Fine-Grained Analysis of the LoRA Fine-Tuning Landscape with Implications for Data Selection

Bowen Zhang, Changrui Fang, Xinsong Ma, Jiaye Teng, Ziye Ma

arXiv 2610.06542首次发表:更新:

发表机构

City University of Hong Kong; Shanghai University of Finance and Economics(香港城市大学; 上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 LoRA-RIP 度量,证明秩过参数化可消除虚假局部极小值,并据此在固定秩预算下指导数据选择,实验验证了秩与数据质量需联合考虑以提升微调效率与可靠性。

AI 中文摘要

低秩适配(LoRA)已成为参数高效微调的标准方法,然而一个基本的实际问题仍未解决:应如何选择适配器秩?过小的秩可能导致优化景观条件不佳,而过大的秩则牺牲了 LoRA 最初追求的微调效率。现有的理论分析对此权衡提供的指导有限,且其保证通常是在严格的理论假设下建立的。我们通过基于非凸低秩矩阵感知的现代结果,为 LoRA 开发了一种显著更精细的景观理论,填补了这一空白。我们的核心见解是,合适的适配器秩应取决于数据诱导的优化几何质量,而非仅取决于模型本身。为形式化这一联系,我们引入了 LoRA-RIP,一种依赖于数据的受限等距度量,用于刻画交叉熵(CE)目标沿 LoRA 相关低秩方向的条件数。我们证明,足够的秩过参数化(所需秩由 LoRA-RIP 常数明确决定)可消除虚假局部极小值,从而将现有基于 RIP 的保证扩展到经典 1/3 区域之外。这一表征还使得在固定秩预算下进行有原则的数据选择成为可能。跨语言和视觉任务的实验支持这些理论预测,表明秩和数据质量是两种耦合资源,应联合考虑以实现更高效、更可靠的 LoRA 微调。

英文摘要

Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the efficiency that motivates LoRA in the first place. Existing theoretical analyses provide only limited guidance on this trade-off, and their guarantees are typically established under restrictive theoretical settings. We address this gap by developing a substantially sharper landscape theory for LoRA, building on modern results from nonconvex low-rank matrix sensing. Our central insight is that the appropriate adapter rank should depend on the quality of the data-induced optimization geometry, rather than on the model alone. To formalize this connection, we introduce LoRA-RIP, a data-dependent restricted-isometry metric that characterizes the conditioning of the cross-entropy (CE) objective along LoRA-relevant low-rank directions. We prove that sufficient rank over-parameterization, with the required rank explicitly determined by the LoRA-RIP constant, eliminates spurious local minima, thereby extending existing RIP-based guarantees beyond the classical 1/3 regime. This characterization further enables principled data selection under a fixed rank budget. Experiments across language and vision tasks support these theoretical predictions, showing that rank and data quality are two coupled resources that should be jointly considered for more efficient and reliable LoRA fine-tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑