arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

宽度扩展作为类增量学习的一种方法

Width Expansion as a Method for Class Incremental Learning

A. L. S. Conde, Y. Elkhatib, C. M. Ranieri

arXiv 2609.37702首次发表:更新:

发表机构

São Paulo State University (UNESP); University of Glasgow(圣保罗州立大学; 格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出动态宽度扩展方法,按归一化损失准则增加神经元数量,结合注意力机制,在Split MNIST和Split CIFAR-100上优于固定架构,有效缓解类增量学习中的灾难性遗忘。

AI 中文摘要

类增量学习(Class-IL)要求模型随时间学习新类别,同时在没有过去数据或任务身份的情况下保留先前获得的知识。这种设置加剧了稳定性-可塑性困境,并使灾难性遗忘成为核心挑战。现有方法包括正则化、知识蒸馏、回放和架构扩展。然而,许多扩展方法依赖于显式的任务标识符或预定义的生长策略,限制了它们在推理时无法获得任务边界的情况下的适用性。本工作提出了一种动态宽度扩展方法,根据归一化损失准则增加现有层内的神经元数量,无需任务特定信息。还引入了一种具有持久键值存储的注意力机制,以稳定特征表示并减少先前学习类别与新引入类别之间的干扰。该方法在标准Class-IL协议下于Split MNIST和Split CIFAR-100上进行了评估。实验比较了固定容量和动态扩展架构,分别有和没有注意力,并结合了包括EWC、LwF和A-GEM在内的既有持续学习方法。结果表明,渐进式宽度扩展始终优于固定架构,特别是在与功能方法和A-GEM结合时。宽度扩展与注意力的组合提供了最一致的增益。总体而言,基于表示需求的动态宽度扩展为Class-IL提供了一种有效且灵活的策略,尽管不受控制的增长可能增加过拟合和计算成本。

英文摘要

Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing approaches include regularization, knowledge distillation, replay, and architectural expansion. However, many expansion methods rely on explicit task identifiers or predefined growth strategies, limiting their applicability when task boundaries are unavailable at inference time. This work proposes a dynamic width expansion method that increases the number of neurons within existing layers according to a normalized loss criterion, without requiring task-specific information. An attention mechanism with persistent key-value memory is also incorporated to stabilize feature representations and reduce interference between previously learned and newly introduced classes. The approach is evaluated on Split MNIST and Split CIFAR-100 under the standard Class-IL protocol. Experiments compare fixed-capacity and dynamically expanding architectures, both with and without attention, combined with established continual learning methods including EWC, LwF, and A-GEM. Results show that progressive width expansion consistently improves performance over fixed architectures, particularly when combined with functional methods and A-GEM. The combination of width expansion and attention provides the most consistent gains. Overall, dynamic width expansion based on representational demand provides an effective and flexible strategy for Class-IL, although uncontrolled growth may increase overfitting and computational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑