发表机构
School of Mathematical Sciences, Peking University; Population Health Sciences Institute, Newcastle University; School of Engineering Mathematics and Technology, University of Bristol; School of Computing and Creative Technologies, University of the West of England(北京大学数学科学学院; 纽卡斯尔大学人口健康科学研究所; 布里斯托尔大学工程数学与技术学院; 西英格兰大学计算与创意技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对任务增量学习中长任务序列下的网络容量问题,提出AdaHAT机制,改进HAT方法,经多数据集实验,其平均性能优于基线,可平衡网络稳定性与可塑性。
AI 中文摘要
灾难性遗忘是任务增量学习中的核心问题,神经网络在训练新任务时易覆盖已学知识。不少基于架构的方法被提出解决该问题,但网络学习长任务序列时,还存在网络容量相关的问题:随着网络在长任务序列中训练越来越多新任务,为防止遗忘,越来越多的活跃参数变为静态。本文提出自适应硬注意力机制(AdaHAT),该机制可结合参数对过往任务的重要性及当前网络容量,对静态参数进行自适应更新。基于此,我们开发了融合AdaHAT机制的新型神经网络架构,AdaHAT在现有基于架构的方法HAT(硬注意力机制)基础上扩展,以更好支持长任务序列的任务增量学习。我们在多个数据集上开展实验,将AdaHAT与包含HAT在内的任务增量学习基线方法对比。实验结果显示,AdaHAT在所有任务上的平均性能优于这些基线,尤其在长任务序列上表现突出,证明了学习此类任务序列时,平衡网络稳定性与可塑性的优势,缓解了网络容量问题。我们的代码可在该http地址获取。
英文摘要
Catastrophic forgetting is a major problem in task-incremental learning, where neural networks tend to overwrite previously learned knowledge when trained on new tasks. A number of architecture-based approaches have been proposed to address this problem. However, the architecture-based approaches suffer from another problem related to network capacity when the networks learn long task sequences: As a network is trained on an increasing number of new tasks in a long task sequence, a growing proportion of active parameters becomes static to prevent forgetting of previously learned knowledge. In this paper, we propose Adaptive Hard Attention to the Task (AdaHAT) with an adaptive attention mechanism which allows adaptive updates to static parameters by taking into account the information about previous tasks on both the importance of these parameters to previous tasks and the current network capacity. Based on this idea, we develop a new neural network architecture incorporating our proposed AdaHAT mechanism. AdaHAT extends an existing architecture-based approach, Hard Attention to the Task (HAT), to better support task-incremental learning over long task sequences. We conduct experiments on a number of datasets and compare AdaHAT with task-incremental learning baselines including HAT. Our experimental results show that AdaHAT achieves better average performance across tasks than these baselines, especially on long task sequences, demonstrating the benefits from balancing the trade-off between stability and plasticity of a network when learning such sequences of tasks, alleviating the network capacity problem. Our code is available at pengxiang-wang.com/projects/continual-learning-arena.
Comments18 pages, 6 figures, published in ECML PKDD 2024
Journal refProceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases. pp. 143-160 (2024)
DOI:10.1007/978-3-031-70352-2_9