arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环神经网络中通过可回收单元门控实现动力系统的持续学习

Continual Learning of Dynamical Systems in Recurrent Neural Networks through Recyclable Unit Gating

Sima Hashemi, Daniel Durstewitz, Georgia Koppe

arXiv 2609.38356首次发表:更新:

发表机构

Heidelberg University; Central Institute of Mental Health (CIMH); University of Tübingen(海德堡大学; 中央精神健康研究所; 蒂宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对循环神经网络持续学习动力系统时的灾难性遗忘问题,提出持续可回收单元门控(CRUG),通过紧凑分配和前向迁移节约容量,实现零遗忘并达到最优重构-容量权衡。

AI 中文摘要

动力系统重构(DSR)旨在从观测时间序列中推断模型,以再现系统的定性长期行为。持续动力系统重构(cDSR)要求在学习新系统的同时保留先前学习的动力学,然而即使循环模型中的微小参数更新也可能在长期自主展开中定性改变其行为。我们在完全可训练且可解释的几乎线性循环神经网络(AL-RNN)上对涵盖参数正则化、重放和参数隔离的既有持续学习(CL)方法进行了基准测试。参数隔离最有效地保留早期动力学,但过度的任务特定分配可能迅速耗尽固定大小的网络。因此,我们引入了持续可回收单元门控(CRUG),通过紧凑分配和前向迁移来节约容量。使用基于$L_0$的惩罚训练的可微门控选择任务特定单元,而未使用的单元则被回收用于后续任务。有向连接允许后续任务重用早期表示,而不影响先前已提交单元的动力学。CRUG在测试方法中实现了最强的重构-容量权衡,零遗忘,并可靠地学习异构的非线性和混沌系统序列。此外,我们表明,当任务共享相似的底层动力学时,前向迁移更为显著和有用。最后,我们证明CRUG的优势超越了自主cDSR,扩展到顺序认知任务。

英文摘要

Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continual DSR (cDSR) requires learning new systems while preserving previously learned dynamics, yet even small parameter updates in recurrent models can qualitatively alter their behavior over long autonomous rollouts. We benchmark established continual learning (CL) methods spanning parameter regularization, replay, and parameter isolation on the fully trainable and interpretable Almost-Linear RNN (AL-RNN). Parameter isolation preserves earlier dynamics most effectively, but excessive task-specific allocations can rapidly exhaust a fixed-size network. We therefore introduce Continually-Recyclable Unit-Gating (CRUG), which conserves capacity through compact allocation and forward transfer. Differentiable gates trained with an $L_0$-based penalty select task-specific units, while unused units are recycled for subsequent tasks. Directed connections allow later tasks to reuse earlier representations without affecting the dynamics of previously committed units. CRUG achieves the strongest reconstruction--capacity trade-off among the tested methods with zero forgetting and reliably learns a heterogeneous sequence of nonlinear and chaotic systems. Furthermore, we show that forward transfer is more pronounced and useful when tasks share similar underlying dynamics. Lastly, we demonstrate that CRUG's advantages extend beyond autonomous cDSR to sequential cognitive tasks.

Comments34 pages (including Appendix), 15 Tables, 7 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑