arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11690cs.LGstat.ML

漂移与依赖:基于重放的持续学习的分层信息论边界

Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning

Tieliang Gong, Zhongbo Zhang, Wen Wen, Yong-Jin Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出分层信息论框架,分解重放式持续学习的泛化差距,通过Wasserstein松弛和SGLD实例化实现可操作的边界,经实验验证了相关预测。

中文摘要 AI 辅助

持续学习需在吸收新任务的同时不遗忘旧任务,而重放——将少量过去样本的缓冲混入当前训练——是缓解灾难性遗忘最有效的方法之一。然而其泛化行为受两个耦合效应影响,现有分析将其归为单一假设级量:有限内存用经验代理替代每个过去分布,重复复用通过共享优化轨迹将缓冲、当前数据和最终假设耦合。我们提出分层信息论框架,在每一层分离这些效应。主要结果将期望泛化差距分解为重放诱导的表示漂移和优化依赖项,后者进一步分解为稳定性、可塑性、交互和残差耦合分量。两项改进使框架可操作:漂移项的Wasserstein松弛(在支持不匹配下有效)产生依赖深度的漂移-敏感性权衡,其最小值确定需稳定的内部层;优化项的SGLD实例将其简化为轨迹级对数行列式预算,揭示感知曲率的梯度对齐统计量,可作为任务级遗忘的在线诊断。受控实验和基准实验证实了预测的内存缩放、内部漏斗以及对齐信号与遗忘的关联。

英文摘要

Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting. Yet its generalization behavior is shaped by two coupled effects that existing analyses fold into a single hypothesis-level quantity: finite memory replaces each past distribution with an empirical proxy, and repeated reuse couples the buffer, the current data, and the final hypothesis through a shared optimization trajectory. We develop a layer-wise information-theoretic framework that separates these effects at every depth. Our main result decomposes the expected generalization gap into a replay-induced representation drift and an optimization-dependence term, the latter further resolved into stability, plasticity, interaction, and residual-coupling components. Two refinements make the framework operational. A Wasserstein relaxation of the drift term, valid under support mismatch, yields a depth-dependent drift--sensitivity trade-off whose minimizer identifies which interior layer to stabilize. An SGLD instantiation of the optimization term reduces it to a trajectory-level log-determinant budget, exposing a curvature-aware gradient-alignment statistic that serves as an online diagnostic of task-wise forgetting. Controlled and benchmark experiments confirm the predicted memory scaling, the interior funnel, and the alignment signal's link to forgetting.

发表机构

  • School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
  • Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)

机构由 AI 辅助整理,请以论文原文为准。

↑