arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线持续学习中的空间一致性可靠重放

Reliable Replay through Spatial Coherence in Online Continual Learning

Haixiang Sun, Jiefu Zhang, Yinghao He, Yang Xu, Vaneet Aggarwal, Bharat Bhargava, Andrew L. Liu

arXiv 2609.33725首次发表:更新:

AI 中文总结

提出SPHERE重放分配方法,利用表示核聚合损失变化并采用熵正则化传输,在多种持续学习任务中提升准确率并减少遗忘。

AI 中文摘要

持续地将模型适应于新任务需要在有限的内存和计算下保留早期知识。经验重放解决了这一挑战,但基于个体损失增加的优先级会忽略相关记忆对同一更新的响应方式,并可能过度强调孤立的响应。我们提出了用于重放的空间一致性风险控制(SPHERE),这是一种通用的重放分配方法,适用于广泛的学习设置。SPHERE使用表示核来聚合带符号的预期损失变化,衰减不支持的尖峰,同时保留一致性的增加。然后,它将分配表述为熵正则化传输,将均匀的源质量重新分配至支持的高风险区域,同时惩罚长距离传输。我们从传输目标对原始损失变化的敏感性中推导出重放系数,并将其与均匀重放混合以维持基线排练。我们的分析建立了核聚合改善风险估计的条件,并限制了由残余噪声和平滑偏差导致的传输价值膨胀。实验表明,SPHERE在噪声标签视觉任务、持续语言模型指令微调以及具有不完整测试奖励的代码生成强化学习中提高了准确率并减少了遗忘。

英文摘要

Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated responses. We introduce SPatial coHErent risk control for REplay (SPHERE), a general replay-allocation method applicable across a broad range of learning settings. SPHERE uses a representation kernel to aggregate signed prospective loss changes, attenuating unsupported spikes while retaining coherent increases. It then formulates allocation as entropy-regularized transport, redistributing uniform source mass toward supported high-risk regions while penalizing long-distance transfers. We derive replay coefficients from the transport objective's sensitivity to the original loss changes and blend them with uniform replay to maintain baseline rehearsal. Our analysis establishes conditions under which kernel aggregation improves risk estimation and bounds transport-value inflation due to residual noise and smoothing bias. Experiments demonstrate that SPHERE improves accuracy and reduces forgetting across noisy-label vision tasks, continual language-model instruction tuning, and code-generation reinforcement learning with incomplete test rewards.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑