arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线持续学习中生长型与弹性型神经网络的可塑性

Plasticity of Growing and Elastic Neural Networks in Online Continual Learning

Jeong Min Kong, Richard S. Sutton

arXiv 2608.01475首次发表:更新:

发表机构

University of California, Los Angeles (UCLA); University of Alberta; Alberta Machine Intelligence Institute (Amii)(加利福尼亚大学洛杉矶分校; 阿尔伯塔大学; 阿尔伯塔机器智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究在线持续学习中生长型与弹性型神经网络的可塑性,实验显示自适应生长型和弹性型神经网络可维持高准确率、不丧失可塑性,或为在线持续学习提供有前景的算法类别。

AI 中文摘要

近年来,可在学习过程中生长或同时生长与收缩的神经网络(分别称为生长型神经网络与弹性型神经网络)已在离线持续学习中得到探索,研究重点聚焦于灾难性遗忘问题。基于以下观察:1)在线持续学习与动物学习的过程高度相似;2)可塑性丧失(即学习网络学习能力的逐步下降)是持续学习面临的另一关键挑战;3)近期研究表明,逐步引入随机初始化的隐藏单元有助于维持可塑性,本文研究了几种基础的生长型与弹性型神经网络在在线持续学习中的可塑性。我们在监督学习场景下开展实验,结果显示:自适应生长型神经网络会逐步向网络中加入新的随机初始化单元,同时保持所有现有连接具有自适应性,尽管死亡隐藏单元的比例持续上升,该网络仍能维持较高的预测准确率且未丧失可塑性。此外,我们还证明,自适应弹性型神经网络除逐步添加新隐藏单元外,还会在每个新任务开始时修剪已估计的死亡隐藏单元,该网络可在维持优异准确率且不丧失可塑性的同时,保持近乎恒定的紧凑规模。我们的结果表明,生长型与弹性型神经网络具备根据相关学习目标调整自身结构的能力,这类算法有望成为在线持续学习中维持高可塑性的有前景的算法类别。

英文摘要

Neural networks that can grow or both grow and shrink during learning, referred to as growing neural networks and elastic neural networks, respectively, have recently been explored in offline continual learning with a particular focus on catastrophic forgetting. Driven by the observations that 1) online continual learning closely resembles how animals learn; 2) loss of plasticity---the progressive decline in a learning network's ability to learn---is another crucial challenge facing continual learning; and 3) incremental introduction of randomly initialized hidden units was recently shown to help preserve plasticity, in this paper, we study the plasticity of several foundational growing and elastic networks in online continual learning. Our experiments in supervised learning settings show that adaptive growing networks, which incrementally incorporate new, randomly initialized units to the network while keeping all existing connections adaptive, can maintain high prediction accuracy without losing plasticity despite the continuous increase in the dead hidden unit proportion. Furthermore, we demonstrate that adaptive elastic networks, which in addition to progressively adding new hidden units also prune estimated dead hidden units at the beginning of each new task, can achieve excellent accuracy without loss of plasticity while simultaneously maintaining a near-constant, compact size. Our results suggest that growing and elastic networks, which exhibit the ability to adapt its structure to the relevant learning objectives, can be a promising class of algorithms also for preserving high plasticity in online continual learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑