arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28184cs.LG

促进多层感知机中Grokking的生物启发机制

Biologically Inspired Mechanisms for Facilitating Grokking in Multilayer Perceptrons

  • Faculty of Automatic Control and Computer Engineering(自动控制与计算机工程学院)
  • “Gheorghe Asachi” Technical University of Iași(雅西“格奥尔基·阿萨奇”技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Florin Leon

AI总结:

该研究探究生物启发机制能否促进多层感知机的Grokking转变,经在稀疏奇偶校验和带噪异或分类基准上实验,发现稳态调节增益最强,为相关机制应用于大语言模型提供支撑。

AI中文摘要:

Grokking是一种从记忆到泛化的延迟转变,常伴随内部表征的重大重组。本文研究:许多未被人工神经网络普遍采用的生物启发机制,能否通过在神经元活动、响应及有效连接层面调控隐藏层计算,主动推动这一转变。我们为多层感知机添加输入门控、结构可塑性、增益调制、阈值调制、稳态调节、侧向抑制及激活去相关机制,并在两个已建立的Grokking基准任务(稀疏奇偶校验和带噪异或分类)上通过系统消融实验评估这些机制。结果显示,各机制对泛化的贡献不均:稳态调节提供最强且最一致的增益,结构稀疏化是第二大重要机制;其余生物启发机制在本次实验中效果较小或一致性较差。对于两个问题,结果支持一项共同原则:明确调控神经元利用率和有效连接可提升可泛化内部计算的出现。这些发现推动对生物启发活动调控及自适应稀疏化的更广泛研究,包括在大语言模型中应用,或可加速可泛化表征的开发并减少实现稳健泛化所需的优化时间。

英文摘要:

Grokking is a delayed transition from memorization to generalization that is often accompanied by substantial reorganization of internal representations. This paper studies whether biologically inspired mechanisms, many of which are not commonly incorporated into artificial neural networks, can actively promote this transition by regulating hidden-layer computation at the levels of neuronal activity, response, and effective connectivity. We augment a multilayer perceptron with input gating, structural plasticity, gain modulation, threshold modulation, homeostasis, lateral inhibition, and activation decorrelation, and evaluate these mechanisms through systematic ablations on two established grokking benchmarks: sparse parity and noisy XOR classification. The results show that the mechanisms contribute unequally to generalization. Homeostasis provides the strongest and most consistent benefit, while structural sparsification emerges as the second major mechanism. The remaining biologically inspired mechanisms have smaller or less consistent effects in the present experiments. For both problems, the results support the common principle that explicit regulation of neuron utilization and effective connectivity can improve the emergence of generalizable internal computation. These findings motivate broader investigation of biologically inspired activity regulation and adaptive sparsification, including in large language models, where they may accelerate the development of generalizable representations and reduce the optimization time required for robust generalization.

补充信息

↑