各向同性曲率下通过联合切空间优化的秩高效LoRA
Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature
浏览论文内容
中文总结 AI 辅助
针对LoRA秩利用率不足问题,提出ISO-LoRA优化器,通过谱下降耦合因子更新,提升有效秩与下游性能,在0.1B-7B模型上验证有效性。
中文摘要 AI 辅助
低秩适配(LoRA)是一种通过学习低秩权重更新来适配大型预训练模型的有效方法。在实践中,LoRA秩用于控制适配器的参数预算和表示能力。我们表明这种观点是不完整的:虽然名义秩决定了表示能力,但优化器塑造了在诱导的权重空间更新中使用多少这种能力。在对GPT-2进行LoRA适配的案例研究中,我们观察到强烈的秩相关优化器效应。尽管使用相同的名义秩,AdamW通常产生具有集中奇异谱和低有效秩的每步更新,而Muon使用更丰富的方向集,并且更一致地从增加LoRA秩中受益。这些观察促使我们提出ISO-LoRA,一种通过权重空间中诱导的切向扰动的谱下降来耦合LoRA因子更新的优化器。ISO-LoRA促进更新在奇异方向上更均匀地分布能量,提高秩利用率,同时保持与LoRA参数化的兼容性。我们通过理论保证补充了这一设计,表明在风格化的尖峰梯度模型下,通过一步分析,ISO-LoRA可以实现比标准因子级优化器更高的有效秩。我们在0.1B-7B参数模型的语言模型适配中验证了这一设计,其中ISO-LoRA提高了有效秩和下游性能,在中等到大的LoRA秩下增益最强。我们的结果强调了秩利用率是LoRA优化的关键因素,并表明优化器设计为更强的参数高效适配提供了重要途径。
英文摘要
Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapter's parameter budget and representational capacity. We show that this view is incomplete: while the nominal rank determines the representational capacity, the optimizer shapes how much of that capacity is used in the induced weight-space updates. In a case study of GPT-2 adaptation with LoRA, we observe a strong rank-dependent optimizer effect. Despite using the same nominal rank, AdamW often produces per-step updates with concentrated singular spectra and low effective rank, whereas Muon uses a richer set of directions and benefits more consistently from increasing LoRA rank. These observations motivate ISO-LoRA, an optimizer that couples the LoRA factor updates through spectral descent on the induced tangent perturbation in weight space. ISO-LoRA promotes updates that distribute energy more evenly across singular directions, improving rank utilization while preserving compatibility with the LoRA parameterization. We complement this design with theoretical guarantees showing that ISO-LoRA can achieve higher effective rank than standard factor-wise optimizers through a one-step analysis under a stylized spiked-gradient model. We validate this design on language-model adaptation across 0.1B-7B-parameter models, where ISO-LoRA improves effective rank and downstream performance, with the strongest gains at moderate-to-large LoRA ranks. Our results highlight rank utilization as a key factor in LoRA optimization and suggest that optimizer design offers an important path toward stronger parameter-efficient adaptation.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
- DRW Associates LLC(DRW联合有限责任公司)
- University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。