arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeuMoSync:面向持续学习中可塑性与适应性的端到端神经调节控制

NeuMoSync: End-to-End Neuromodulatory Control for Plasticity and Adaptability in Continual Learning

Seyed Roozbeh Razavi Rohani, Khashayar Khajavi, Wesley Chung, Mandana Samiei, Mo Chen

arXiv 2608.04358首次发表:更新:

发表机构

Simon Fraser University; Mila - Quebec AI Institute(西蒙菲莎大学; 米拉-魁北克人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出新型架构 NeuMoSync,借鉴大脑神经调节机制,为深度神经网络添加神经元特异性调节模块,在多种持续学习基准上提升了模型的可塑性与前后向适应能力。

AI 中文摘要

持续学习(CL)要求模型按顺序学习任务,但深度神经网络常面临可塑性丧失和知识迁移效果差的问题,这会阻碍其长期适应性。受大脑全局神经调节机制的高层次启发,我们提出了神经调节与同步(NeuMoSync),这是一种将动态、神经元特异性调节集成到深度神经网络中的新型架构,以增强模型的适应性和可塑性。NeuMoSync 为每个神经元扩展了标准神经网络架构,配备可学习的特征向量,用于跟踪网络范围的历史上下文,同时配备一个处于更高抽象层级的模块。该模块基于当前输入和网络的演化状态,合成神经元特异性信号,以自适应调节激活动力学和突触可塑性。在多种持续学习基准上进行评估,包括记忆任务(随机标签 CIFAR-10 和随机标签 MNIST)、概念漂移任务(打乱 CIFAR-10 和打乱 Mini-ImageNet)、类增量学习任务(类拆分 ImageNet 和类拆分 CIFAR-100)以及域增量学习任务(置换 MNIST),NeuMoSync 在保持可塑性方面表现出强大性能,且与现有方法相比,在前向和后向适应上均取得了提升。消融研究验证了每个组件的必要性,对学习到的调节信号的分析揭示了跨任务的可解释协调模式。我们的研究强调了将全局协调机制集成到深度学习系统中,以推进鲁棒、自适应的持续学习的潜力。代码可在指定 URL 公开获取。

英文摘要

Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor knowledge transfer, which can impede their long-term adaptability. Drawing high-level inspiration from global neuromodulatory mechanisms in the brain, we introduce Neuromodulation and Synchronization (NeuMoSync), a novel architecture that integrates dynamic, neuron-specific modulation into deep neural networks to enhance their adaptability and plasticity. NeuMoSync extends standard neural network architectures with learnable feature vectors for each neuron that track network-wide historical context and with a module operating at a higher level of abstraction. This module synthesizes neuron-specific signals, conditioned on both current inputs and the network's evolving state, to adaptively regulate activation dynamics and synaptic plasticity. Evaluated on diverse CL benchmarks, including memorization (Random Label CIFAR-10 and Random Label MNIST), concept drift (Shuffle CIFAR-10 and Shuffle Mini-ImageNet), class-incremental learning (Class Split ImageNet and Class Split CIFAR-100), and domain-incremental learning (Permuted MNIST), NeuMoSync demonstrates strong performance in retaining plasticity and achieves improvements in both forward and backward adaptation compared with existing methods. Ablation studies validate the necessity of each component, while analysis of the learned modulatory signals reveals interpretable coordination patterns across tasks. Our work underscores the potential of integrating global coordination mechanisms into deep learning systems to advance robust, adaptive continual learning. The code is publicly available at https://github.com/RoozbehRazavi/NeuMoSync.

CommentsPublished in Transactions on Machine Learning Research (TMLR)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑