AI 中文总结
本研究通过受控顺序训练实验证明,神经网络在行为收敛后仍保留训练历史导致的表征差异,表明行为收敛不蕴含表征收敛。
AI 中文摘要
朝向相同最终目标训练的神经网络可以达到相似的预测性能,同时保留由早期训练历史塑造的内部表征。我们通过受控的顺序训练实验研究这一效应,实验中配对的卷积网络从相同的权重出发,经历相反的任务顺序,然后接收相同的确定性共同松弛分布。在20次配对的MNIST运行中,16次满足预先声明的行为匹配标准,但其匹配的表征保留了平均历史分数0.139(95%自助法置信区间:0.127-0.153)和约3.1%的预测分歧。将共同松弛扩展到50,000次优化器更新并未消除测量到的差异:在五个配对种子中,表征历史分数在测量时间范围结束时仍为0.190(95%自助法置信区间:0.161-0.219),而平均准确率差距仅为0.18个百分点。新的线性探针表明,在具有足够标注数据的情况下,两种历史保留了实际上等价的线性可访问类别信息。一个相同标签的旋转MNIST对照实验重现了该效应:所有五个配对种子均达到行为匹配,同时保留平均表征历史分数0.162。最后,一个匹配学习率的ReLU-LeakyReLU对照实验在五个配对种子中平均将50,000次更新的表征残留减少了0.040,提供了方向性证据表明激活介导的可塑性有助于训练历史效应的持续存在。这些结果提供了协议范围内的证据,表明行为收敛不一定意味着表征收敛,且优化历史在长时间共同训练后可以留下可测量的内部痕迹。
英文摘要
Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by earlier training history. We study this effect using controlled sequential-training experiments in which paired convolutional networks start from identical weights, experience reversed task orders, and then receive the same deterministic common-relaxation distribution. Across 20 paired MNIST runs, 16 satisfy a predeclared behavioral-matching criterion, yet their matched representations retain a mean history score of 0.139 (95% bootstrap CI: 0.127-0.153) and approximately 3.1% prediction disagreement. Extending common relaxation to 50,000 optimizer updates does not erase the measured difference: across five paired seeds, the representation-history score remains 0.190 (95% bootstrap CI: 0.161-0.219) at the end of the measured horizon while the mean accuracy gap is only 0.18 percentage points. Fresh linear probes show that, with sufficient labeled data, the two histories retain practically equivalent linearly accessible class information. A same-label rotated-MNIST control reproduces the effect: all five paired seeds reach behavioral matching while retaining a mean representation-history score of 0.162. Finally, a matched-learning-rate ReLU-LeakyReLU control reduces the 50,000-update representation residue by 0.040 on average in all five paired seeds, providing directional evidence that activation-mediated plasticity contributes to the persistence of training-history effects. These results provide protocol-scoped evidence that behavioral convergence need not imply representational convergence and that optimization history can leave measurable internal traces after prolonged common training.
Comments12 pages, 6 figures, 2 tables. Code and reproducibility artifacts: https://github.com/Ertugrulmutlu/hysteresis-neural-networks