发表机构
The University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究LoRA微调中秩与双下降的关系,发现秩为1时风险最低,存在第二次下降但需超越参数对等才能匹配密集微调性能。
AI 中文摘要
双下降现象引起了广泛关注,近期工作将其与数据、模型及学习配置联系起来。实际微调通常涉及在冻结的预训练权重之上训练一个小型适配器,如低秩适配(LoRA)。适配器的秩是设定其容量的超参数,然而该秩与双下降的关系尚未得到充分探索。我们在标签噪声下,对四个视觉骨干网络和一个7B语言模型进行了量化,采用模块匹配秩扫描(MMRS),该方法扩展至超过全秩,并将每个秩与相同模块的密集微调(按种子配对)进行比较。在DeiT-Tiny上,风险在秩为1时最低,随着适配器能够拟合噪声标签而急剧上升,形成插值悬崖。越过峰值后,风险再次下降,但每个测试的峰值后秩(仍节省参数)均保持在密集风险之上。秩为1的LoRA在四个视觉骨干网络中的三个上优于密集微调,这与小秩时的强正则化一致。因此,LoRA表现出第二次下降,但仅在失去其参数优势后才匹配密集风险,首先是在DeiT-Tiny上,其投影权重为密集的四倍。代码可在https URL获取。
英文摘要
Double descent has sparked considerable interest, with recent work relating it to the data, the model and the learning configuration. Practical fine-tuning commonly involves training a small adapter on top of frozen pretrained weights, as in low-rank adaptation (LoRA). The adapter's rank is the hyperparameter that sets its capacity, yet how this rank relates to double descent has not been well explored. We quantify this relation under label noise on four vision backbones and a 7B language model with a module-matched rank sweep (MMRS), which extends past full rank and compares every rank with dense fine-tuning of the same modules, paired by seed. On DeiT-Tiny, risk is lowest at rank one and rises sharply as the adapter becomes able to fit the noisy labels, forming an interpolation cliff. Past the peak, risk falls again, but every tested post-peak rank that still saves parameters remains above dense risk. Rank-one LoRA outperforms dense fine-tuning on three of the four vision backbones, consistent with strong regularization at small rank. LoRA thus exhibits a second descent, but matches dense risk only after losing its parameter advantage, first on DeiT-Tiny at four times dense's projection weights. Code is available at https://anonymous.4open.science/r/lora-second-descent-C7B5.
Comments30 pages, 14 figures, 15 tables