学习轨迹上的损失参数化费希尔宽度
Loss-Parameterized Fisher Width Along Learning Trajectories
- FPT University(FPT大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究探究学习轨迹上费希尔宽度的演化,推导相关分解与界,在特定模型中证明总体梯度流选择的分支满足极限关系,实验验证分支式损失参数化费希尔宽度的合理性。
AI中文摘要:
费希尔宽度用于衡量探测集经局部费希尔几何变形后的高斯宽度,本文研究其沿学习轨迹的演化,并探究训练损失能否作为该量的有效坐标。首先推导了针对固定紧探测集的精确迹-形状分解与确定性稳定性界;在总体高斯教师逻辑模型中,教师对齐状态在低于log2的每个损失水平上均为极值:其参数范数最小,且同时最大化费希尔迹与欧氏球费希尔宽度。随后证明总体梯度流渐近选择该分支,对齐与正交坐标具有显式速率,对d≥2,有w_F(B_2^d;θ(t))/√L(θ(t))→√6/π E[χ_{d-1}]。受控全费希尔实验验证了匹配损失分支与总体预测;在采用对角模型-费希尔近似的非线性多层感知机(MLP)中,梯度下降(GD)与随机梯度下降(SGD)在匹配损失处仍接近,而Adam则遵循显著偏移的分支,所测试的固定探测集保持高度相似的时间形状。这些结果支持费希尔宽度采用分支式而非通用的损失参数化方式。
英文摘要:
Fisher width measures the Gaussian width of a probe set after deformation by the local Fisher geometry. We study its evolution along learning trajectories and ask when training loss can serve as an effective coordinate for this quantity. We first derive an exact trace--shape factorization and a deterministic stability bound for fixed compact probes. In a population Gaussian-teacher logistic model, the teacher-aligned state is extremal on every loss level below $\log 2$: it has minimal parameter norm and maximizes both Fisher trace and Euclidean-ball Fisher width. We then show that population gradient flow asymptotically selects this branch, with explicit rates for the aligned and orthogonal coordinates. This yields, for $d\geq2$, \[ \frac{w_F(B_2^d;θ(t))} {\sqrt{L(θ(t))}} \longrightarrow \frac{\sqrt6}π\mathbb E[χ_{d-1}]. \] Controlled full-Fisher experiments support the matched-loss branch and the population predictions. In a nonlinear MLP with a diagonal model-Fisher approximation, GD and SGD remain close at matched loss, whereas Adam follows a substantially displaced branch; the fixed probes tested retain highly similar temporal shapes. These results support a branchwise, rather than universal, loss parametrization of Fisher width.