联合仿射谱塑形:超越仅权重Muon的权重与偏置更新耦合
Control Allocation in Neural Network Optimization: Joint Affine Control of Weight and Bias Updates
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对IMDb数据集训练BERT-mini,对比不同谱优化方法,提出的联合正则化逆方法可提升测试准确率、降低损失,是仅权重谱优化的小幅但一致的扩展。
AI中文摘要:
矩阵谱优化器会重塑权重更新谱,但通常将向量值偏置交由单独的优化器处理。本研究探讨这种分离是否中立。我们将每个仿射层建模为联合动量矩阵A=[M_W,α m_b],并对完整矩阵应用带上限的正则化逆谱映射,同时生成权重和物理偏置的更新。我们在从零开始训练的四层BERT-mini模型上,针对IMDb数据集开展严格的五组随机种子消融实验,对比了精确SVD Muon、仅权重逆塑形、仿射探测逆塑形,以及本文提出的联合正则化逆(JRI)方法。仅权重逆塑形将验证损失选定的测试准确率从84.903±0.242%提升至85.562±0.308%,并将选定的测试损失从0.3479降至0.3345。允许偏置在保留独立Adam偏置更新的同时改变联合SVD,并未比仅权重逆塑形表现更好。联合使用变换后的偏置将选定的测试准确率提升至85.738±0.180%,测试损失降至0.3291,且五组随机种子的结果均优于探测基线。在峰值性能窗口内,JRI在维持合格权重更新范数的同时,将偏置更新范数从0.02095降至0.00301,边界函数占比从86.58%降至78.97%,并使权重诱导的边界运动与显式偏置之间的余弦值从+0.030变为-0.137。独立的22组随机种子重复实验得出选定测试准确率为85.743±0.203%。这些结果表明,联合仿射谱分配是对仅权重谱优化的小幅但一致的扩展。
英文摘要:
Optimization algorithms determine not only the magnitude of a neural-network update but also how that update is distributed across parameter channels. We study whether this distribution can be treated as a controllable quantity independently of global training progress. We define operational update allocation through normalized channel energies and analyze two scalar controls: a coordinate-preconditioning exponent and an affine spectral exponent that scales the bias column of an augmented weight--bias matrix. At a frozen state, a common nonzero step-size multiplier leaves normalized allocation unchanged; the coordinate exponent yields affine pairwise log-odds with an explicit inverse; and the affine exponent induces a rank-one positive-semidefinite Gram perturbation and a logistic raw-participation law. We further separate raw affine participation, spectral gain, and the decoded physical bias update, and show that finite polynomial spectral iterations preserve singular subspaces. Same-state replay verifies the exact control laws. On a five-seed controlled benchmark, intermediate controls improve held-out and worst-group metrics, whereas excessive affine control causes underfitting. A four-task single-seed transfer study provides descriptive corroboration. These results establish instantaneous allocation control and a bounded empirical operating regime, but do not imply a task-independent generalization ordering.