arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

持续学习机制组合实现长时程记忆

Continual Learning Mechanisms Compose for Long-Horizon Memorization

Zheyuan Zhang, Alvin Zhang, Daniel Khashabi, Tianmin Shu

arXiv 2609.06986首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语言模型长时程记忆中的灾难性遗忘,提出组合数据、函数和权重锚点及合并LoRA的持续学习方法,在100任务上平均保留率提升28倍。

AI 中文摘要

语言模型可能需要内化随时间到达的信息,并在后续多次更新中保留这些信息。为了研究这一挑战,我们引入了长时程记忆(long-horizon memorization)这一设置,在该设置中,模型通过持续监督微调学习100个查询-答案任务,而不保留早期训练示例或在推理时接收任务标识符。顺序更新会导致灾难性遗忘,在我们评估的持续学习机制中,没有单一机制能够在此时间跨度上保持较强的记忆保持能力。我们假设,针对遗忘互补来源的机制在组合时会更有效。我们沿着两个设计维度组织这些组合。数据、函数和权重锚点(anchors)指定每次更新应保留的先前信息,而低秩分配规则决定后续更新在何处保留。为系统验证这一假设,我们构建了三个不同的100任务记忆数据集。我们引入任务级连续减半(task-level successive halving)来搜索组合设计空间,并使用因子实验测量个体效应和交互效应。我们最好的方法将三种锚点与合并的LoRA(merged LoRA)相结合,在所有数据集中排名前3,并将平均最终保留率从朴素顺序微调下的1.2%提高到34.9%,实现了28倍的改进。数据锚点和合并LoRA提供了最大的平均增益,并在所有三个数据集上表现出超加性交互。综合这些结果,表明组合互补机制显著改善了长时程记忆,超越了任何单一机制所能达到的效果。

英文摘要

Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training examples or receiving task identifiers at inference. Sequential updates cause catastrophic forgetting, and no single continual learning mechanism we evaluate maintains strong retention at this horizon. We hypothesize that mechanisms addressing complementary sources of forgetting will be more effective when composed. We organize these compositions along two design dimensions. Data, function, and weight anchors specify what prior information each update should preserve, while low-rank allocation rules determine where successive updates are retained. To test this hypothesis systematically, we construct three distinct 100-task memorization datasets. We introduce task-level successive halving to search the combinatorial design space and use a factorial experiment to measure individual and interaction effects. Our best method combines all three anchors with merged LoRA, ranks among the top 3 methods in all datasets, and raises average final retention from 1.2% under naive sequential fine-tuning to 34.9%, a 28-fold improvement. The data anchor and merged LoRA provide the largest average gains and interact super-additively on all three datasets. Together, these results show that composing complementary mechanisms substantially improves long-horizon memorization beyond what any individual mechanism achieves.

CommentsProject page: https://compose-cl.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑