arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19995cs.NE

前向-反向脱节:神经计算中的状态动力学、信用分配与生物学基础

The Forward-Backward Disconnect: State Dynamics, Credit Assignment, and Biological Grounding in Neural Computation

Hadi Al Mubasher, Mariette Awad

AI总结:

本文指出神经计算中前向计算已多样化但训练方法仍集中于梯度类误差传播机制的前向-反向脱节问题,提出涵盖状态动力学等三要素的分类法,强调需对齐三者以弥合该脱节。

AI中文摘要:

神经计算中反复出现的一种模式是,将动力学和生物学结构重新引入原本为可扩展优化而简化的模型中。早期的前馈网络将生物神经元简化为阈值或类速率求和单元,这种抽象与大规模全局梯度训练兼容。此后,前向计算已多样化:现代架构携带循环状态,通过注意力和关联记忆从长上下文检索信息,通过结构化状态空间动力学压缩历史,在连续时间上演化,收敛到隐式平衡点,并通过稀疏尖峰通信。训练的多样化程度较低,可扩展学习仍集中在反向传播、随时间反向传播、伴随方法、隐式微分和替代梯度变体上。我们将这种不对称性称为“前向-反向脱节”,并沿三个耦合轴开发了涵盖神经模型家族的分类法:状态动力学结构、信用分配机制和生物学基础。前向和学习基础被分开处理,分析单位是架构-学习配置,而非仅架构名称。在静态、循环、基于注意力、状态空间、连续时间、隐式、尖峰、生物学上合理和神经形态家族中,前向动力学已多样化,但已证明的最高规模仍集中在全局或紧密衍生自梯度的误差传播机制中。弥合这一脱节需要状态动力学、信用分配和计算基底之间更好的对齐。

英文摘要:

A recurring pattern in neural computation is the reintroduction of dynamical and biological structure into models originally simplified for scalable optimization. Early feedforward networks reduced biological neurons to threshold or rate-like summation units, an abstraction compatible with global-gradient training at scale. Since then, forward computation has diversified: modern architectures carry recurrent state, retrieve from long contexts through attention and associative memory, compress histories through structured state-space dynamics, evolve in continuous time, settle to implicit equilibria, and communicate through sparse spikes. Training has diversified less. Scalable learning remains concentrated around backpropagation, backpropagation through time, adjoint methods, implicit differentiation, and surrogate-gradient variants. We call this asymmetry the forward-backward disconnect and develop a taxonomy spanning neural model families along three coupled axes: state-dynamics structure, credit-assignment mechanism, and biological grounding. Forward and learning grounding are treated separately, and the unit of analysis is the architecture-learning configuration rather than the architecture name alone. Across static, recurrent, attention-based, state-space, continuous-time, implicit, spiking, biologically plausible, and neuromorphic families, forward dynamics have diversified while the highest demonstrated scales remain concentrated in global or closely gradient-derived error-propagation mechanisms. Closing this disconnect requires better alignment among state dynamics, credit assignment, and computational substrate.

↑