arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

参数高效微调方法真的不同吗?

Are Parameter-Efficient Fine-tuning Methods Really Different?

Yikuan Li, Pinyan Lu, Fanghui Liu

arXiv 2610.09122首次发表:更新:

发表机构

The University of Warwick; Shanghai University of Finance and Economics; Shanghai Jiao Tong University(华威大学; 上海财经大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究比较六种参数高效微调方法,发现频谱保持并非必要,LoRA限制遗忘、DoRA性能更优,建议通过功能后果评估几何约束。

AI 中文摘要

参数高效微调(PEFT)提供了多种参数化方式,但它们在方法论和功能上的差异仍不清楚。我们比较了语言模型和扩散模型中的六种方法,以考察其参数化方式与任务性能、遗忘以及预训练权重几何结构变化之间的关系。受正交微调(OFT)的频谱保持设计启发,我们首先探究频谱保持本身对适应和保留是否重要。我们发现,所选的LoRA系列方法也近似保持了预训练几何结构,并且恢复它们轻微漂移的奇异值谱在很大程度上保持了任务性能,这质疑了显式几何保持的必要性。此外,我们观察到一些方法表现出不同的适应-保留权衡,且这些权衡在不同设置下有所变化:LoRA在竞争性性能下最一致地限制了遗忘,DoRA在大多数比较中比LoRA取得了更高的平均任务分数,而PiSSA往往带来更大的保留成本。进一步的干预实验表明,虽然不同PEFT方法的性能提升可归因于对不同频谱组件组的修改,但我们一致发现,恢复主导组件而非中间或尾部组件,会在通用文本NLL或基础图像漂移方面产生最大的平均减少。总之,这些结果促使我们通过功能后果而非单纯保持来评估几何约束。代码可在以下网址获取:此 https URL。

英文摘要

Parameter-efficient fine-tuning (PEFT) offers many parameterizations, yet their methodological and functional differences remain unclear. We compare six methods in language and diffusion models to examine how their parameterizations relate to task performance, forgetting, and changes in pretrained weight geometry. Motivated by the spectrum-preserving design of orthogonal fine-tuning (OFT), we first ask whether spectral preservation is itself important for adaptation and retention. We find that the selected LoRA-family methods also approximately preserve pretrained geometry, and that restoring their slightly drifted singular-value spectra largely preserves task performance, questioning the necessity of explicit geometric preservation. Beyond this, we observe that some methods exhibit distinct adaptation--retention trade-offs that vary across settings: LoRA most consistently limits forgetting at competitive performance, DoRA achieves higher mean task scores than LoRA in most comparisons, while PiSSA often incurs greater retention costs. Further intervention experiments suggest that while performance gains from different PEFT methods can be attributed to modifications in different groups of spectral components, we consistently find that restoring dominant rather than intermediate or trailing components produces the largest mean reduction in general-text NLL or base-image drift. Together, these results motivate evaluating geometric constraints through their functional consequences rather than preservation alone. Code is available at https://github.com/Kuaaannn/PEFT_methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑