发表机构
School of Mathematical Sciences, Shanghai Jiao Tong University; Institute of Natural Sciences, MOE-LSC, Shanghai Jiao Tong University; Huawei Technologies Ltd; Shanghai University of Finance and Economics(上海交通大学数学科学学院; 上海交通大学自然科学研究院; 华为技术有限公司; 上海财经大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对两层网络的模加法学习,提出概率签名方法揭示梯度训练选择傅里叶回路的机制,还解释了标签噪声下的损失变化现象,且可推广至异或等其他算子。
AI 中文摘要
在模加法任务上训练的神经网络通常会形成支持精确泛化的傅里叶结构表征。尽管已有研究识别出这些傅里叶回路,但基于梯度的训练如何从数据分布中选择它们的机制仍不清楚。我们利用概率签名解决该问题,概率签名通过训练分布的条件统计量表达主导梯度相互作用。对于模加法,这些签名是循环移位算子,经离散傅里叶变换对角化后产生近似解耦的傅里叶模态动力学,这解释了傅里叶稀疏性、频率匹配和相位对齐的出现。同一框架解决了标签噪声下的一个谜题:尽管损坏的样本缺乏连贯的泛化规则,却能比干净样本表现出更快的早期损失下降。我们表明,噪声会增加条件标签冲突,强化早期共享坐标的增强。最后,该方法可应用于其他算子,以异或(XOR)为例,我们在实验中观察到了预测的频率。
英文摘要
Neural networks trained on modular addition tasks often develop Fourier-structured representations that support exact generalization. While prior work has identified these Fourier circuits, the mechanism by which gradient-based training selects them from the data distribution remains unclear. We address this question using probability signatures, which express leading gradient interactions through conditional statistics of the training distribution. For modular addition, these signatures are cyclic shift operators and are diagonalized by the discrete Fourier transform, yielding approximately decoupled Fourier-mode dynamics. This explains the emergence of Fourier sparsity, frequency matching, and phase alignment. The same framework resolves a puzzle under label noise: corrupted examples can show faster early loss decrease than clean examples, despite lacking a coherent generalization rule. We show that noise increases conditional label collisions, strengthening early shared-coordinate reinforcement. Finally, this method can be applied to other operators. Taking XOR as an example, we observed the predicted frequency in experiments.
Comments36 pages, 18 figures. Submitted to ICLR 2027