arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32124cs.LG

我们应该冻结什么?受保护的冻结:连通性塑造预训练模型的微调

What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models

Leonel Aguilar

首次发表
浏览论文内容

中文总结 AI 辅助

针对微调中冻结权重可能失效的问题,提出基于连通性的受保护冻结策略,通过 removal-value 和 drift-value 选择冻结层,在多个模型上提升旧任务保留准确率。

中文摘要 AI 辅助

在通过微调适配预训练模型时,仅冻结权重可能无法保持性能,因为其他位置的更新会改变冻结核心的输入,最终影响整体性能。我们首先分析所选冻结核心可以被隔离的情况,并提出 removal-value,一种基于容量的评分,近似 HOPE 的移除代价在移除顺序上的平均值。我们表明,在 VGG-8 中,切断从可训练神经元到冻结核心的路径使得使用该评分的选择有效:冻结 70% 时,旧任务准确率比 DEFT 高 5.22±0.51 个百分点,而新任务准确率相近。在 Transformer 中,共享残差流使得到冻结神经元的路径保持开放。针对这种情况,我们推导出 drift-value,一种前向传播的代理,用于估计在局部更新模型下更新每个权重条目引起的输出扰动。在语言模型中,在 40 个 epoch 时,该策略在具有显著保留损失的情况下超过适配的 Wanda 和 RIA 冻结评分,而其与 Fisher 的差异仍未解决。在 Qwen2.5-1.5B 上经过 160 个 epoch 后,它比静态 Fisher 多保留 0.0433±0.0102。在 DINOv3 视觉 Transformer 适配点云时,drift-value 保留 0.440 的图像准确率,而相同数量的随机掩码为 0.187。这些结果促使了受保护的冻结:当传入路径被切断时,根据 removal-value 选择;当路径保留时,根据 drift-value 选择。

英文摘要

When adapting pre-trained models through fine-tuning, freezing weights alone might not preserve performance, as updates elsewhere can change the inputs to the frozen core, ultimately affecting overall performance. We first analyse the case where a selected frozen core can be isolated and propose removal-value, a capacity-based score that approximates HOPE's removal cost averaged over removal orders. We show that in VGG-8, cutting paths from trainable neurons into a frozen core makes selection using this score useful: $70\%$ frozen preserves $5.22\pm0.51$ percentage points more old-task accuracy than DEFT at similar new-task accuracy. In transformers, shared residual streams leave paths into frozen neurons open. For this case, we derive drift-value, a forward-only proxy for the output disturbance from updating each weight entry under a local update model. In language models, at 40 epochs, this policy exceeds adapted Wanda and RIA freezing scores in settings with substantial retention loss, while its differences from Fisher remain unresolved. After 160 epochs on Qwen2.5-1.5B, it retains $0.0433\pm0.0102$ more than static Fisher. In DINOv3 vision-transformer adaptation to point clouds, drift-value retains $0.440$ image accuracy versus $0.187$ for a random mask of the same count. These results motivate Guarded Freezing: select by removal-value when incoming paths are cut, and by drift-value when they remain.

发表机构

  • ETH Zürich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑