arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当预条件指数变为负值时:学习率耦合与跨环境泛化

When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization

Gongyue Zhang, Honghai Liu

arXiv 2609.30271首次发表:更新:

AI 中文总结

本研究通过四环境分类实验发现,最大化跨环境泛化的预条件指数随学习率对数线性下降,负指数并非普遍最优,而是高学习率下的分配机制,并揭示了源域验证与跨环境鲁棒性之间的模型选择冲突。

AI 中文摘要

自适应优化器通常由二阶矩估计的固定幂次参数化。现有的部分自适应方法研究了介于动量式更新与标准Adam平方根之间的指数,而该指数与全局学习率之间的相互作用则较少被理解。我们使用一个配对的四环境分类问题进行了受控的跨环境研究,该问题包含稳定的稀疏特征、环境依赖的虚假稀疏特征、稠密特征以及高维噪声。在覆盖21个预条件指数 $p\in[-0.5,0.5]$ 和五个学习率 $\eta\in[10^{-4},10^{-2}]$ 的 \NumRuns{} 次源训练运行中,我们发现最大化跨环境准确率的指数几乎随 $\log_{10}\eta$ 线性下降。拟合斜率范围从 $-0.270$ 到 $-0.300$,$R^2$ 介于 $0.972$ 和 $0.996$ 之间。在 $\eta=10^{-2}$ 时,源验证选择在所有四个环境中仍偏好正指数,而跨环境和最差环境准则则偏好负指数。检查点分解表明,较低的 $p$ 降低了学习到的虚假-稳定和噪声-稳定权重比;在相关性反转下,它还降低了有害虚假边际的幅度。因此,负 $p$ 并非普遍最优设置。它是由学习率和预条件共同作用产生的高步长分配机制。该研究还揭示了模型选择冲突:源域验证系统性地选择了与最大化对环境变化鲁棒性不同的预条件机制。结果是一项单种子、有限预算的机制研究,而非广泛的基准声明。

英文摘要

Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and the global learning rate is less understood. We perform a controlled cross-environment study using a paired four-environment classification problem with stable sparse features, environment-dependent spurious sparse features, dense features, and high-dimensional noise. Across \NumRuns{} source-training runs covering 21 preconditioning exponents $p\in[-0.5,0.5]$ and five learning rates $η\in[10^{-4},10^{-2}]$, we find that the exponent maximizing cross-environment accuracy decreases almost linearly with $\log_{10}η$. The fitted slopes range from $-0.270$ to $-0.300$, with $R^2$ between $0.972$ and $0.996$. At $η=10^{-2}$, source-validation selection still prefers positive exponents in all four environments, whereas cross-environment and worst-environment criteria prefer negative exponents. Checkpoint decomposition shows that lower $p$ reduces the learned spurious-to-stable and noise-to-stable weight ratios; under reversed correlation, it also reduces the magnitude of the harmful spurious margin. Negative $p$ is therefore not a universally optimal setting. It is a high-step-size allocation regime produced by the joint action of learning rate and preconditioning. The study also exposes a model-selection conflict: source-domain validation systematically selects a different preconditioning regime from the one that maximizes robustness to environmental change. The results are a single-seed, finite-budget mechanism study rather than a broad benchmark claim.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑