发表机构
Department of Artificial Intelligence and Data Science, Sri Ramakrishna Engineering College, Anna University; Interdisciplinary Institute for Applied AI and Data Science Ruhr (AKIS), Department of Electrical Engineering and Computer Science, Bochum University of Applied Sciences(人工智能与数据科学系,斯里·拉玛克里希纳工程学院,安娜大学; 鲁尔应用人工智能与数据科学跨学科研究所(AKIS),电气工程与计算机科学系,波鸿应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用阿米尔乔回溯线搜索获取损失函数局部锐度信息,通过接受步长估算方向曲率,给出低成本在线稳定性边界读数,还能给出学习率上限保障Adam对大初始学习率的鲁棒性。
AI 中文摘要
损失函数的局部锐度,即顶部海森矩阵特征值λ1,决定了最大稳定梯度步长,但通常测量它需要兰佐斯或海森向量迭代。我们观察到单次阿米尔乔回溯线搜索已经以几次前向传递的成本携带了此信息:接受步长α在回溯因子设置的乘法带内包围了方向曲率q = g⊤Hg/||g||²。在CIFAR-10、Fashion-MNIST和Imagenette上,logα跟踪logλ1,皮尔逊相关系数为-0.91至-0.95,给出低成本在线稳定性边界读数。在初始化时使用一次此测量可产生学习率上限(一种保障,而非更快的优化器),使Adam在超过三个数量级(10⁻³至3.0)的范围内对过大的初始学习率具有鲁棒性,开销约为1%,并且当所选速率已经安全时它不执行任何操作。一次探测就足够了:训练期间的定期探测不会带来更强的鲁棒性益处。原始梯度探测揭示了机制,但需要通过一分钟的发散扫描针对架构校准安全因子。沿Adam自身更新方向进行探测消除了这种校准:单个固定安全因子κ = 2可避免我们测试的所有九个架构以及所有四个基准的完整学习率网格上的发散,并且该方法可直接应用于AdamW。
英文摘要
Local sharpness, defined by the largest Hessian eigenvalue $λ_1$, sets the maximum stable gradient update size, but its computation would usually require running Lanczos or Hessian-vector products. However, we notice that even a single Armijo backtracking line search already contains this information with just a few forward passes, as the accepted step $α$ determines the directional curvature along the search direction up to the multiplicative band set by the backtracking factor. The correlation between $\logα$ and $\logλ_1$ on CIFAR-10, Fashion-MNIST and Imagenette reaches $-0.91$ to $-0.95$ in Pearson correlation, and even after removing the trend per run the correlation remains at $-0.60$ to $-0.70$. This allows for a cheap online Edge-of-Stability estimate of the slow sharpness component. The employed probing mechanism searches along Adam's first-step update direction, at initialisation and nine times over the course of the first 50 optimiser steps. The learning-rate cap is set as twice the smallest observed step size, and in the studied learning-rate ranges ($10^{-3}$ to $3.0$) and GPT-2 pretraining experiments all capped runs avoid divergence; in the general architecture analysis, one MLP architecture is still sensitive to the initial batch order. This probing protocol incurs an approximately one percent overhead, and using a non-binding cap means that the optimiser's state and first update are bit-identical. There is no fine-tuning of any of the protocol parameters to any specific architecture; this is the sense in which this is a calibration-free safeguard. It is meant as a way of avoiding divergence, not achieving accuracy. At GPT-2 scale, multiple measurements also illustrate why an initialisation-only cap is insufficient: directional curvature rises substantially in the first five optimiser steps, which motivates the short probationary window.
Comments38 pages, 7 figures. Published in Transactions on Machine Learning Research (10/2026): https://openreview.net/pdf?id=foLGwG3ne4. Code: https://github.com/jfrochte/armijo-curvature-probe
Journal refTransactions on Machine Learning Research (ISSN 2835-8856), October 2026