arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01947q-bio.NCcs.AIcs.NE

归一化分裂塑造连续工作记忆的低秩慢流形

Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

  • School of System Science, Beijing Normal University(北京师范大学系统科学学院)
  • State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(北京师范大学认知神经科学与学习国家重点实验室)
  • Qiyuan Laboratory(启元实验室)

机构由 AI 辅助整理,请以论文原文为准。

Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang

AI总结:

本文提出循环归一化分裂网络(RDNN),通过动力学分析和BPTT梯度分析,证明归一化分裂可使网络形成稳健低秩慢流形,防止时变输入下的流形破碎,是学习高保真连续表示的关键机制。

AI中文摘要:

稳健维持与更新连续变量的能力是工作记忆的核心特征。经典连续吸引子网络存在严重的微调脆弱性,而GRU、LSTM等标准人工循环神经网络(RNN)通常无法稳定学习连续流形,反而会将状态空间破碎为离散点吸引子。为弥合这一差距,受皮质回路中广泛存在的经典神经计算——归一化分裂(divisive normalization)启发,本文提出循环归一化分裂网络(Recurrent Divisive Normalization Network, RDNN),这是一种最小且代数孤立的动态除法模型。通过对经典工作记忆任务的动力学系统分析,我们证明这种生物物理约束可使网络收敛到稳健、高保真的慢流形。此外,我们分析了随时间反向传播(Backpropagation Through Time, BPTT)过程中归一化分裂的梯度动力学,表明其引入了依赖活动的局部梯度缩放,该缩放会在高活动状态下抑制参数更新,经验上与网络有效秩的显著自压缩一致,将循环动力学限制在紧凑的低维子空间,同时避免了与显式低秩分解相关的优化问题。最后,消融实验显示,尽管减法抑制可维持静态记忆,但归一化分裂在数学上对防止时变输入下的流形破碎至关重要。我们的研究结果表明,归一化分裂不仅是生物的附属产物,更是学习高保真连续表示的关键计算机制。

英文摘要:

The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor networks suffer from severe fine-tuning fragility, standard artificial recurrent neural networks (RNNs) like GRUs and LSTMs typically fail to stably learn continuous manifolds, instead shattering the state space into discretized point attractors. To bridge this gap, we draw inspiration from divisive normalization, a canonical neural computation widely observed across cortical circuits, and propose the Recurrent Divisive Normalization Network (RDNN), a minimal and algebraically isolated model of dynamic division. Through dynamical systems analysis on canonical working memory tasks, we demonstrate that this biophysical constraint allows the network to converge to robust, high-fidelity slow manifolds. Furthermore, we analyze the gradient dynamics of divisive normalization during Backpropagation Through Time (BPTT), showing that it introduces an activity-dependent local gradient scaling. This scaling dampens parameter updates in highly active regimes, which empirically aligns with a significant self-compression of the network's effective rank, confining the recurrent dynamics to a tight, low-dimensional subspace while avoiding the optimization pathologies associated with explicit low-rank factorization. Finally, ablations demonstrate that while subtractive inhibition can maintain static memories, divisive normalization is mathematically essential to prevent manifold shattering under time-varying inputs. Our findings identify divisive normalization not merely as a biological artifact, but as a critical computational mechanism for learning high-fidelity continuous representations.

补充信息

↑