等变性打破学习率
Equivariance Breaks the Learning Rate
- University of Stuttgart(斯图加特大学)
- Bitdefender(比特 Defender(比特防御者))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文发现等变网络中不可约表示块的梯度结构导致Adam学习率不匹配,提出块归一化方法,结合动量调整使Adam在势模型上媲美Muon。
AI中文摘要:
等变网络通常使用 Adam 优化器进行训练,然而近期研究报道,诸如 Muon 等矩阵结构优化器在这些架构上可以表现更好,但未解释原因。我们在等变线性层内部识别出这一差异的一个来源。每个不可约表示块学习一个通道混合矩阵 $W_l$,该矩阵在其 $2l+1$ 个分量之间共享,从而得到扩展映射 $W_l \otimes I_{2l+1}$。对于该层的单次应用,$W_l$ 的梯度累加 $2l+1$ 个外积贡献,其秩至多为 $2l+1$。Adam 在不使用不可约表示边界的情况下单独重新缩放存储的权重,因此一个学习率可以在层内不同块之间产生不同的谱步长。我们通过分别归一化每个块的更新来解决这种不匹配,而不引入新的超参数。这仅改变更新的尺度,保持 Adam 的矩估计及其在每个块内的方向不变。我们在一个受控的 $\mathrm{SO}(3)$ 等变模型(带有匹配的稠密对照)以及一个在 rMD17 和 MD22 上训练的 e3nn 原子间势模型中评估该机制。玩具设置隔离了一种随宽度增长的不匹配,而稠密对照未显示相应的增长。在原子间势模型中,块归一化和调整 Adam 的动量系数独立地提升性能,但两者单独均无法匹敌 Muon。两者结合使 Adam 在所有数据集上与 Muon 具有竞争力,表明块级步长控制和动量累积解释了 Muon 的大部分优势。
英文摘要:
Equivariant networks are commonly trained with Adam, yet recent work reports that matrix structured optimizers such as Muon can perform better, with the reasons for these gains only partly understood. We identify one source of this difference inside equivariant layers. An equivariant layer learns one channel mixing matrix $W_l$ per degree $l$, which we call an irrep block, and shares it across the $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. This sharing sums gradient contributions across components and can produce different update scales under SGD. Adam's entrywise normalization reduces sensitivity to gradient scale, but neither optimizer directly controls the effective step size of each block. A single learning rate can therefore produce different effective step sizes across blocks. Muon instead controls the effective step size by approximately equalizing the singular values of each momentum matrix. We normalize each irrep block update by a single scalar, preserving its singular value ratios while letting the learning rate control its size. We implement this with spectral normalization or a simpler root-mean-square normalization. We evaluate spectral normalization in a controlled $\mathrm{SO}(3)$-equivariant model with a matched non-equivariant model. In this setting, the step size mismatch grows with width in the equivariant model but not in the non-equivariant model. We evaluate both variants across molecular force prediction on the rMD17 and MD22 datasets, QM9 molecular property prediction, and charged particle dynamics. Across these applications, block normalization generally improves Adam and closes part of its gap to Muon. These results highlight an overlooked interaction between equivariant architectures and their optimizers. Studying and designing the two together may help explain and address training difficulties often attributed to equivariance itself.