arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05402cs.LGcs.CL

任务向量下降:从非独立同分布批次中学习

Task Vector Descent: Learning from Non-IID Batches

发表机构慕尼黑工业大学 · 苏黎世联邦理工学院
查看机构详情
  • Technical University of Munich(慕尼黑工业大学)
  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Anton Baumann, Jonas Hübotter, Zeynep Akata, Andreas Krause

首次发表
浏览论文内容

中文总结 AI 辅助

针对非独立同分布批次导致的稳定性-可塑性权衡,提出任务向量缩放方法,通过部分应用参数位移提升持续学习性能,优于完全整合。

中文摘要 AI 辅助

持续学习中的一个核心挑战是在获取新知识的同时不遗忘模型已学到的内容。当训练数据来自不同领域、用户或任务特定的分布,且这些分布随时间不均匀地出现时,这一挑战在语言模型训练中尤为明显。在此类设置中,连续的迷你批次按时间聚类于特定分布,而非从整体数据混合中独立同分布采样。对时间聚类数据进行训练会引发稳定性-可塑性权衡。使模型适应当前活跃分布可以提升模型在该分布上的表现,但可能导致其在先前训练数据上的性能下降。我们发现,随着对同一分布的暴露时间延长,这种权衡会加剧。因此,我们探究由这种序列产生的参数位移(即任务向量)应完全保留还是部分应用。我们将完全应用位移(λ=1)与部分整合进行比较,后者在将任务向量应用于持续模型之前先按λ缩放,并以相同系数缩放优化器状态。在持续预训练、从随机初始化预训练、监督后训练和强化后训练中,我们发现中间λ值通常能提升持续模型的平均性能,相对于完全整合而言,尤其是在较长的同分布序列之后。在受控数据流和自然定义的适应序列的持续预训练实验中,任务向量缩放以匹配的学习率优于完全整合,表明其益处不能仅通过学习率缩放来复现。

英文摘要

A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, successive minibatches are temporally clustered by distribution instead of being sampled i.i.d. from the overall data mixture. Training on temporally clustered data induces a stability-plasticity tradeoff. Adapting the model to the active distribution can improve the model on the active distribution but may lead to a performance degradation on data it previously trained on. We find that this tradeoff intensifies with longer exposure to the same distribution. We therefore ask if the parameter displacement produced by such a sequence (the task vector) should be fully retained or applied partially. We compare applying the full displacement ($λ=1$) with partial integration, which scales the task vector by $λ$ before applying it to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, we find that intermediate values of $λ$ often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In continual-pretraining experiments with both controlled streams and naturally defined adaptation sequences, task-vector scaling outperforms full integration at the matched learning rate, showing that its benefits are not reproduced by learning-rate scaling alone.

↑