发表机构
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Fisher8是一种输出层梯度修正方法,通过Fisher几何稳定神经异方差回归,无数据依赖超参数,在多维回归等任务中实现了更优的似然-误差权衡与校准不确定性估计。
AI 中文摘要
从含噪观测中训练神经网络以联合预测均值和不确定性估计的过程存在不稳定性,催生了一系列独立的稳定化研究。我们指出,这些干预措施凸显了一个共性底层问题:梯度步长与损失景观的几何结构匹配度较差。为使更新与局部曲率更好匹配,我们推导了Fisher8,这是一种输出层梯度修正方法,它利用Fisher几何而非欧氏几何来重新定向和缩放更新。与以往稳定器不同,Fisher8除学习率外无数据依赖的超参数,且能在连续预测分布间引入近似KL信任半径。我们证明,现有稳定器收敛于该几何修正的重叠分量。在多维回归和表征学习任务中,Fisher8实现了更优的似然-误差权衡,预测出校准后的不确定性估计,并学习到丰富的感知不确定性特征空间。
英文摘要
Training neural networks to jointly predict mean and uncertainty estimates from noisy observations can be unstable, prompting a series of independent stabilization efforts. We argue that these interventions highlight a common underlying issue where gradient steps are poorly aligned with the geometry of the loss landscape. To better align updates with local curvature, we derive Fisher8, an output-layer gradient correction that reorients and rescales updates using Fisher geometry rather than Euclidean geometry. Unlike past stabilizers, Fisher8 introduces no data-dependent hyperparameters beyond learning rate and admits an approximate KL trust radius between successive predictive distributions. We show that prior stabilizers converge on overlapping components of this geometric correction. Across multidimensional regression and representation-learning tasks, Fisher8 obtains superior likelihood--error tradeoffs, predicts calibrated uncertainty estimates, and learns rich uncertainty-aware feature spaces.