κ-LoRA:条件数揭示哪些LoRA矩阵值得更新
\k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating
浏览论文内容
中文总结 AI 辅助
研究指出LoRA统一更新矩阵计算成本高,条件数大的矩阵对性能提升贡献大。提出κ-LoRA方法,聚焦更新条件数大的矩阵,可减半可训练参数数量,降低计算和内存成本,实验证明其能缩短微调时间、降低内存成本且不影响精度。
中文摘要 AI 辅助
低秩自适应(LoRA)已成为神经网络高效微调的广泛采用技术,将模型更新分解为低秩矩阵。然而,LoRA计算成本高,因为它统一更新所有矩阵,而不考虑其对自适应的实际贡献。对于数十亿参数的大规模模型以及边缘部署和设备上微调等资源受限设置,此成本尤其高昂。我们首次表明并非所有LoRA矩阵都同样值得调整:条件数较小的矩阵在各方向上已平衡良好,对自适应贡献小;条件数大的矩阵包含欠发达方向,跨越更丰富子空间并推动大部分性能提升。基于此,我们提出κ-LoRA,通过将更新聚焦于条件数最大的矩阵来优化LoRA。通过将LoRA更新限制在按条件数排名前50%的权重矩阵,κ-LoRA将可训练参数数量减半,相应降低计算和内存成本。多个基准测试的广泛实验表明,该设计平均将微调时间缩短16.2%,同时匹配标准LoRA的精度并将内存成本降低4.5%。进一步分析表明所选矩阵的条件数在训练过程中持续下降,这表明κ-LoRA的有效性源于有针对性的谱重新平衡而非仅参数选择。
英文摘要
Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices uniformly, regardless of their actual contribution to adaptation. This cost is especially prohibitive for large-scale models with billions of parameters and for resource-constrained settings such as edge deployment and on-device fine-tuning. We show for the first time that not all LoRA matrices are equally worth tuning: matrices with smaller condition numbers (the ratio of largest to smallest singular value) are already well-balanced across directions and contribute only marginally to adaptation, whereas matrices with larger condition numbers contain underdeveloped directions that span richer subspaces and drive most of the performance gains. This observation itself is a key contribution of our work, and it motivates a more selective approach to fine-tuning. Building on this insight, we propose \k{appa}-LoRA, a method that optimizes LoRA by focusing updates on the matrices with the largest condition numbers, which capture the most informative directions of change. By restricting LoRA updates to the top 50% of weight matrices ranked by condition number, \k{appa}-LoRA halves the trainable parameter count and correspondingly reduces compute and memory cost. Extensive experiments across multiple benchmarks show that this design cuts fine-tuning time by 16.2% on average while matching the accuracy of standard LoRA and reducing memory cost by 4.5%. Further analysis reveals that the condition numbers of the selected matrices consistently decrease over training, suggesting that \k{appa}-LoRA's effectiveness stems from targeted spectral rebalancing rather than parameter selection alone.
发表机构
- King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。