学习重塑内部表征中的幂律各向异性
Learning reshapes power-law anisotropy in internal representations
浏览论文内容
中文总结 AI 辅助
该研究通过求解线性神经网络学习动力学,揭示输入统计与任务结构的动态相互作用是幂律内部表征的形成机制,且非线性网络中存在类似指数演化规律。
中文摘要 AI 辅助
在从最先进的语言模型到小鼠大脑皮层的多种生物和人工神经系统中,均观察到内部表征存在幂律各向异性。这种各向异性是高维信息处理的关键几何属性,是多种理论分析的基础。然而,它如何从输入结构和任务驱动学习中产生的机制仍不清楚。本文中,我们在具有幂律输入和教师结构的教师-学生设置中,通过精确求解宽两层线性神经网络的学习动力学来表征该形成过程。我们表明,在特征学习 regime 中,内部表征谱的局部幂律指数在训练过程中呈非单调演化,且在模式和训练时间上表现出多达四种不同的渐近 regime;相比之下,在懒惰 regime 中,指数基本保持不变。我们进一步通过数值证明,更现实的非线性网络中会出现类似的指数动力学。综合来看,这些结果揭示了一种通用机制:输入统计与任务结构之间的动态相互作用产生了幂律内部表征。
英文摘要
Power-law anisotropy in internal representations has been observed across a wide range of biological and artificial neural systems, from state-of-the-art language models to the mouse cerebral cortex. This anisotropy is a key geometric property of high-dimensional information processing and underlies a variety of theoretical analyses. However, the mechanism by which it emerges from input structure and task-driven learning has remained unclear. Here, we characterize this formation process by exactly solving the learning dynamics of a wide two-layer linear neural network in a teacher--student setting with power-law input and teacher structures. We show that, in the feature-learning regime, the local power-law exponent of the internal-representation spectrum evolves nonmonotonically over the course of training and exhibits up to four distinct asymptotic regimes across modes and training times. By contrast, in the lazy regime, the exponent remains essentially unchanged. We further demonstrate numerically that similar exponent dynamics arise in more realistic nonlinear networks. Together, these results suggest a general mechanism by which the dynamic interaction between input statistics and task structure gives rise to power-law internal representations.
发表机构
- Graduate School of Informatics(情报学研究科)
- Kyoto University(京都大学)
机构由 AI 辅助整理,请以论文原文为准。