arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15854cs.LGcs.CV

遗忘的几何:持续学习中的表示通量

Geometry of Forgetting: Representation Flux in Continual Learning

Maksim A. Kazanskii

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出表示通量作为遗忘的几何标记,提出FlowLess-R表示空间正则化方法,在多个基准测试中结合现有重放方法提升了持续学习的平均准确率并减少了灾难性遗忘。

中文摘要 AI 辅助

灾难性遗忘仍是持续学习的核心障碍,即神经网络在学习新任务时会丢失先前获得的知识。现有方法主要通过参数正则化或经验重放来缓解遗忘,但与遗忘相关的表示空间动力学仍鲜为人知。我们研究序列学习过程中的潜在表示演化,并提出表示通量(representation flux),这是一种衡量训练过程中样本级表示位移的几何指标。我们在多个基准测试中表明,表示通量与灾难性遗忘密切相关,时间分析显示,升高的通量可先于后续性能下降出现。表示位移还与置信度下降相关,而互补的几何属性能提供关于样本级遗忘的额外信息。基于这些观察,我们提出FlowLess-R,一种表示空间正则化方法,该方法相对于存储的参考来约束重放表示,同时允许继续学习。FlowLess-R与架构无关,可通过表示匹配项集成到基于重放的方法中。在SplitMNIST、SplitFashionMNIST、SplitCIFAR10和SplitTinyImageNet上的实验表明,其结合ER、DER++和ER-ACE可提高最终平均准确率并减少遗忘。我们的结果确定表示通量是一种有用的遗忘几何标记,并表明稳定潜在表示提供了一种缓解灾难性遗忘的简单策略。

英文摘要

Catastrophic forgetting remains a fundamental obstacle to continual learning, where neural networks lose previously acquired knowledge while learning new tasks. Existing methods primarily mitigate forgetting through parameter regularization or experience replay, while the representation-space dynamics associated with forgetting remain less understood. We investigate latent representation evolution during sequential learning and introduce representation flux, a geometric measure of sample-level representation displacement across training. We show that representation flux is strongly associated with catastrophic forgetting across multiple benchmarks, with temporal analyses indicating that elevated flux can precede subsequent performance degradation. Representation displacement is also associated with confidence degradation, while complementary geometric properties provide additional information about sample-level forgetting. Motivated by these observations, we propose FlowLess-R, a representation-space regularization method that constrains replay representations relative to stored references while allowing continued learning. FlowLess-R is architecture-agnostic and integrates into replay-based methods through a representation-matching term. Experiments on SplitMNIST, SplitFashionMNIST, SplitCIFAR10, and SplitTinyImageNet show improved final average accuracy and reduced forgetting with ER, DER++, and ER-ACE. Our results identify representation flux as an informative geometric marker of forgetting and show that stabilizing latent representations provides a simple strategy for mitigating catastrophic forgetting.

↑