arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00403cs.LG

逻辑与记忆的冲突:在浅层MLP中学习高阶交互

The Conflict Between Logic and Memory: Training Conditions for Optimizer-Dependent Rule Acquisition

Gongyue Zhang, Honghai Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过合成任务考察浅层MLP在拟合训练数据与学习高阶规则之间的分离,发现优化器与干扰权重学习显著影响高阶交互规则的实现。

中文摘要 AI 辅助

一个网络可以拟合其训练样本,却无法恢复生成这些样本标签的规则。我们通过使用控制交互阶数和干扰输入存在的合成任务,在单隐藏层多层感知机(MLPs)中考察这种分离现象。我们建立了基本的基准性质:纯奇偶校验不包含可预测的低阶边际,允许精确的贝叶斯后验,并且可以在干净潜在输入上由宽度为$k$的ReLU网络表示。实验随后识别出不同的优化结果。在匹配的阶数2-4扫描中,SGD、Adam和Muon在二阶时均达到100%的峰值测试准确率;在三阶时,它们分别达到96.25%、50.87%和76.82%,而Muon在四阶时达到99.21%。在另一个混合阶数任务中,仅冻结与独立干扰输入连接的第一层权重,将AdamW在第10个epoch的准确率从44.73%提高到95.07%。仅在测试时移除相同输入则将其提高到48.38%。因此,干扰权重学习改变了训练结果,其影响超出了对预测的直接效应。偏差干预揭示了目标对称性与浅层ReLU表示之间的联系。在紧凑的仅信号机制中,SGD和Muon都能学习五阶到八阶,而SGD在九阶到十一阶具有更高的峰值准确率。总之,这些结果展示了优化和干扰学习如何约束浅层网络所实现的高阶规则。

英文摘要

Optimizers can fit the same task while acquiring different generalizing relations. We study the training conditions governing these differences in single-hidden-layer ReLU networks, combining composite evidence tasks, parameter-level interventions, and a three-seed strict-parity scan. Our central finding is that nuisance-connected trainability reshapes both shared failures and relative optimizer advantages. In a nuisance-heavy task, all twenty tested optimizer configurations remain near chance on the hardest stage. Retaining every input but fixing nuisance-connected first-layer weights at initialization raises that stage's accuracy from approximately 50\% to 70.56\%, 68.47\%, and 69.00\% for momentum SGD, Adam, and Muon. Masking the same inputs only after full training does not recover this performance. On a separate pairwise-mode task, background freezing reduces Muon's rare-mode advantage over momentum SGD by 12.48 percentage points, while the target and mode frequencies remain fixed. Each intervention is evaluated under a common validation-selection protocol with condition-specific learning rates and checkpoints. A strict-parity sweep over orders 1--20 provides a complementary reference without spurious cues or extra nuisance coordinates: the optimizers separate at orders 9--11, then approach chance despite substantial remaining Bayes predictability. A mixed task establishes a recovery boundary, and CIFAR-10 supplies an external architecture comparison. Together, these findings connect optimizer comparison to the acquisition and use of specified relations, identifying permitted adaptation as a concrete training variable that changes what a fixed architecture learns.

↑