AI 中文总结
本文探究神经网络训练中是否存在趋同进化,提出基于匹配的MLP结构权重比较框架,发现同一任务训练的网络在结构空间中形成任务相关吸引子,且早期学习以微妙分布式调整为主。
AI 中文摘要
在进化生物学中,无亲缘关系的生物在面临相似功能需求时,可独立进化出相似结构。本文探究神经网络训练过程中是否会发生类似形式的趋同进化:具有不同随机初始化的网络,在同一任务上训练时是否会形成相似的内部权重结构?该问题具有技术挑战性,因为隐藏神经元可被任意置换而不改变所表示的函数,这使得直接的矩阵比较具有误导性。我们引入一种基于匹配的框架,用于在结构权重空间中比较多层感知器(MLP)。首先使用置换不变特征对隐藏神经元进行粗略对齐,再通过迭代匈牙利匹配进行精细调整。对齐后,采用专门设计的结构距离度量比较网络,以突出与任务相关的权重模式。将该方法应用于在MNIST、Fashion-MNIST和KMNIST上训练的小型MLP集成,我们发现,在同一任务上训练的网络之间的距离,比与不同任务上训练的网络之间的距离更近。因此,特定任务的训练似乎会引导初始随机网络走向结构网络空间中的不同区域或吸引子。训练的最早阶段揭示了一种额外的意外现象:分类准确率在匹配的结构距离显示出强任务特异性分离、以及全局权重分布发生明显变化之前就已快速上升,但单个权重条目已开始协同漂移。这表明早期学习可能首先通过微妙的分布式调整进行,这些调整对函数影响显著,却几乎不改变网络的粗粒度形态。我们将这种早期形态发生视为未来将研究的更丰富动态过程的初步线索。
英文摘要
In evolutionary biology, unrelated organisms can independently evolve similar structures when exposed to similar functional demands. Here we ask whether an analogous form of convergent evolution occurs during neural network training: do networks with different random initializations develop similar internal weight structures when trained on the same task? This question is technically nontrivial because hidden neurons can be arbitrarily permuted without changing the represented function, making direct matrix comparisons misleading. We introduce a matching-based framework for comparing multilayer perceptrons in structural weight space. Hidden neurons are first coarsely aligned using permutation-invariant features and then refined by iterative Hungarian matching. After alignment, networks are compared with structural distance metrics designed to emphasize task-relevant weight patterns. Applying this approach to ensembles of small MLPs trained on MNIST, Fashion-MNIST, and KMNIST, we find that networks trained on the same task remain closer to one another than to networks trained on different tasks. Thus, task-specific training appears to guide initially random networks toward distinct regions, or attractors, in structural network space. The earliest phase of training reveals an additional and unexpected phenomenon. Classification accuracy rises rapidly before the matched structural distances show strong task-specific separation, and before the global weight distribution visibly changes. Nevertheless, individual weight entries already begin to drift in a coordinated manner. This suggests that early learning may first operate through subtle, distributed adjustments that strongly affect function while leaving coarse network morphology almost unchanged. We treat this early morphogenesis as a first glimpse of a richer dynamical process that will be investigated in future work.