arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00771cs.LG

记忆任务之间迁移的定位

Localizing Transfer Between Memorization Tasks

Yimiao Yu, Florentin Guth

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过随机映射记忆任务间的迁移实验,发现等价与非等价两种迁移模式,并定位其分别源于最后一层的幅度驱动和其他层的结构驱动效应,深化了对迁移学习机制的理解。

中文摘要 AI 辅助

迁移学习中的一个核心谜团是,为什么在一个任务上进行预训练能够加速另一个任务的训练或提升其性能,以及这种迁移背后的机制是什么。在本工作中,我们研究了随机输入-输出映射的记忆任务之间的迁移。我们发现了两种令人惊讶的迁移模式:等价迁移,即每增加一个预训练周期大约节省一个下游微调周期;以及非等价迁移,即在不匹配任务上的预训练甚至可能比直接在下游任务上训练更高效。通过消融实验,我们将迁移分解并定位为两个独立的效果:一个位于最后一层的“平凡”的由幅度驱动的迁移,以及一个“非平凡”的由结构驱动的迁移,部分可归因于其他层的协方差。这些结果加深了我们对迁移学习底层机制的理解,并有可能带来有原则的预训练策略。

英文摘要

A central puzzle in transfer learning is why pre-training on one task can accelerate training or improve performance on another task, and what mechanisms underlie this transfer. In this work, we examine the transfer between memorization tasks of random input-output mappings. We find two surprising transfer patterns: equivalent transfer, where each additional pre-training epoch saves approximately one downstream fine-tuning epoch; and non-equivalent transfer, where pre-training on a mismatched task can be even more efficient than directly training on the downstream task itself. Through ablation experiments, we decompose and localize the transfer into two separate effects: a "trivial" magnitude-driven transfer in the last layer, and a "non-trivial" structure-driven transfer, partially attributable to the covariance of the other layers. These results advance our understanding of the underlying mechanisms of transfer learning and have the potential to lead to principled pre-training strategies.

补充信息

↑