发表机构
I-X Centre for AI in Science, Imperial College London; Department of Electrical and Electronic Engineering, Imperial College London(AI科学中心,帝国理工学院伦敦分校; 电气与电子工程系,帝国理工学院伦敦分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨人工神经网络中低维表征对泛化的作用,通过信息瓶颈使循环神经网络学习低维表征,借助信息论度量刻画其动态,发现非单调轨迹,小鼠实验也有类似结果,表明学习表征对泛化有功能优势及因果作用。
AI 中文摘要
降维已被证明在识别神经流形方面很强大,神经流形是高维神经活动背后的低维结构,改善了群体水平编码的可解释性。但低维表征在学习系统中是否具有生物学相关性并赋予功能优势,还是仅反映神经元水平活动,在神经科学中仍存在争议。我们表明,在时间序列预测任务中,迫使循环神经网络学习低维表征的显式信息瓶颈对于旋转和分布外泛化是必要的。使用因果涌现的信息论度量,我们刻画了这种表征在记忆到泛化转变过程中的动态,发现了一条非单调轨迹。对小鼠学习交替迷宫任务时CA1海马体活动的分析揭示了类似的跟踪行为表现的非单调涌现动态。这些发现表明神经网络学习紧凑、分布式和涌现表征的能力赋予了泛化的功能优势,支持了学习表征在认知中的因果作用。
英文摘要
Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional structures underlying high-dimensional neural activity. These low-dimensional representations have improved the interpretability of population-level coding. Yet whether such low-dimensional representations are biologically relevant and confer functional advantages in learning systems, or merely reflect neuron-level activity, remains contested in neuroscience. We show that an explicit information bottleneck forcing a recurrent neural network to learn a low-dimensional representation is necessary for rotational and out-of-distribution generalisation in a time-series prediction task. Using information-theoretic measures of causal emergence, we characterise the dynamics of this representation across the memorisation-to-generalisation transition, finding a non-monotonic trajectory which shows an initial decrease, a minimum, and a subsequent rise to a maximum, even as prediction loss falls monotonically. This trajectory scales with task complexity, and the magnitude of emergent structure reliably predicts generalisation performance. Analysis of CA1 hippocampal activity in mice learning an alternating maze task reveals analogous non-monotonic emergence dynamics that track behavioural performance. Together, these findings indicate that the ability of neural networks to learn compact, distributed and emergent representations confers a functional advantage for generalisation, supporting a causal role for learned representations in cognition.