隐藏的尺度参数控制ReLU网络中的特征特化
Hidden Gauge Controls Feature Specialization in ReLU Networks
浏览论文内容
中文总结 AI 辅助
该研究在高斯教师-学生模型中发现,初始预测器不可见的隐藏尺度参数可控制ReLU网络中神经元的特征所有权,实现确定性的特征选择与冗余神经元剪枝。
中文摘要 AI 辅助
训练会改变神经网络的预测结果,同时在其内部单元间分配与任务相关的结构。在过参数化的ReLU网络中,多个神经元初始时具有完全相同的功能角色,但其中一个可能会获得教师特征,而其他神经元则变得冗余。我们将该神经元的身份称为特征所有权,并探究其是否可被初始预测器不可见的参数选择所控制。在一个易处理的高斯教师-学生模型中,我们固定完整的初始函数,仅改变正齐次缩放尺度参数。相反的尺度参数会产生不同的特征轨迹,且在特化时间上存在尖锐的Θ(D²)分离,这是全局时钟变化无法解释的。在任意数量的初始重复学生神经元中,将有利的尺度参数分配给一个神经元会确定性地选择它作为所有权所有者,并使其余神经元的功能贡献降至零。精确的反应-传输分解将该效应归因于改变特征系数和方向的不同迁移率。我们证明了全局选择和功能剪枝,将有限时间选择扩展到可见扰动和小步长全批量梯度下降,并在总体和有限样本训练中验证了预测的损失、对齐、剪枝和耗散轨迹。因此,初始预测器既不决定何时学习特征,也不决定哪个神经元学习该特征。
英文摘要
The success of deep learning depends on learning useful representations, yet predicting how training organizes these representations across neurons remains difficult. In this work, we show that changing the scale of initial weights can determine which neurons learn a feature without altering any neuron's initial contribution. We construct ReLU networks with identical initial features and predictions that reach the same final predictions with different roles for their neurons. In one, all neurons share the learned feature. In the other, one neuron acquires it while every other neuron's contribution vanishes. The only change is the relative scale of each neuron's input and output weights. Our analysis explains how an initial learning advantage persists through convergence: as one neuron learns the target, it reduces the error driving the others and limits their subsequent adaptation. We prove this outcome in a nonlinear model under gradient flow and small-step gradient descent, and quantify how scale changes the speed and path of feature learning. Experiments verify the predicted dynamics and show that scale also changes feature assignment when two features compete. These results reveal how initialization can control the organization of a learned representation without changing what the network initially represents.
发表机构
- School of Future Technology Southeast University(东南大学未来技术学院)
机构由 AI 辅助整理,请以论文原文为准。