arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

适配何处至关重要:用于能力保留的层选择性微调

Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention

Zhiqiang Pang, Zihong Sun, Qi Xie, Jun Shu, Deyu Meng, Zongben Xu

arXiv 2610.11620首次发表:更新:

发表机构

School of Mathematics and Statistics, Xi’an Jiaotong University(西安交通大学数学与统计学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对参数高效微调中目标任务适配与通用能力保留的权衡问题,提出层选择性LoRA(LS-LoRA)方法,仅在输入-输出相似度低的层放置LoRA适配器,在数学推理和代码生成任务上实现了目标任务性能提升与通用能力保留的平衡。

AI 中文摘要

参数高效微调(PEFT)使大语言模型(LLMs)能够适配特定任务,但往往以牺牲预训练期间获得的通用能力为代价。现有方法主要通过数据重放或正则化来缓解这种权衡,依赖额外数据或显式优化约束。我们转而关注一个不同的问题:应在何处应用适配?我们发现,对不同Transformer层进行微调会产生不同的目标任务增益和能力退化程度,这表明并非所有层都同样适合适配。为表征这种差异,我们使用分层经验Fisher信息来测量目标任务敏感性。然而,计算Fisher分数需要反向传播计算,对于大型模型而言成本越来越高。因此,我们引入输入-输出余弦相似度作为轻量级、仅前向传播的代理,用于对层敏感性进行排名。在不同模型和任务中,输入-输出相似度较低的层始终表现出更高的经验Fisher分数。基于这一观察,我们提出层选择性LoRA(LS-LoRA),仅在输入-输出相似度低的层中放置可训练的LoRA适配器。在数学推理和代码生成任务上的实验表明,与标准的全层LoRA相比,LS-LoRA提高了平均目标任务性能,同时保留了显著更多的常识推理能力,证明精心选择适配位置可提供一种简单有效的方法来平衡目标任务适配与通用能力保留。

英文摘要

Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining. Existing approaches primarily mitigate this trade-off through data replay or regularization, relying on additional data or explicit optimization constraints. We instead focus on a different question: where should adaptation be applied? We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation. To characterize this difference, we use layer-wise empirical Fisher information to measure target-task sensitivity. However, computing Fisher scores requires backward computation and becomes increasingly expensive for large models. We therefore introduce input--output cosine similarity as a lightweight, forward-only proxy for ranking layer sensitivity. Across models and tasks, layers with lower input--output similarity consistently exhibit higher empirical Fisher scores. Building on this observation, we propose Layer-Selective LoRA (LS-LoRA), which places trainable LoRA adapters only in layers with low input--output similarity. Experiments on mathematical reasoning and code generation show that LS-LoRA improves average target-task performance while retaining substantially more commonsense reasoning capability than standard all-layer LoRA, demonstrating that carefully choosing where to adapt can provide a simple and effective way to balance target-task adaptation and general capability retention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑