自然梯度下降的效率有多低?从精确最优到 Θ(√log d) 发散
How Inefficient Is Natural Gradient Descent? From Exact Optimality to Θ( \sqrt{ \log d } ) Divergence
浏览论文内容
中文总结 AI 辅助
本文量化自然梯度下降相对 Fisher--Rao 测地线的低效率比 R,证明其从精确最优(二次势族)到随维度 d 按 Θ(√log d) 增长(尺度族乘积)的三种情形,并指出计算代价随 R 增加。
中文摘要 AI 辅助
自然梯度下降(NGD)是机器学习中常见方法的基础。对于对偶平坦族,在正向 Kullback--Leibler 目标上的理想化 NGD 遵循混合测地线,该路径通常比最短的 Fisher--Rao 路径更长。我们通过低效率比 R ≥ 1(混合测地线的 Fisher 长度除以 Fisher--Rao 距离)来量化这一额外开销,并将其在端点对上的上确界界定为参数维度 d 的函数。一个张量准则识别出 (I) 类族,其中 R=1 处处成立:恰好是那些具有二次势或一维的族,例如固定协方差的高斯分布。对于非二次族,我们证明另外两种情形:(II) 有界三阶偏度加上有限 Fisher--Rao 直径产生一个与维度无关的界;(III) 对于尺度族的乘积——包括高斯协方差和 Gamma 率——R 增长为 Θ(√log d),在 d 上无界。在每步 Fisher 弦预算下,R 转化为实际计算成本:NGD 渐近地至少需要 R 倍于遵循 Fisher--Rao 测地线的优化器的步数。实验证实了所有三种情形:对于二次势族 (I),R 达到机器精度;分类界 π/(2√2) 被接近但未达到 (II);采样的尺度乘积 R 随 d 增长,对于长距离高维移动达到 R ≈ 1.5 (III)。
英文摘要
Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this overhead by the inefficiency ratio \(R \ge 1\), the Fisher length of the mixture geodesic divided by the Fisher--Rao distance, and bound its supremum over endpoint pairs as a function of the parameter dimension \(d\). A tensor criterion identifies the regime (I) families, with \(R=1\) everywhere: exactly those with quadratic potential or dimension one, such as fixed-covariance Gaussians. For non-quadratic families, we prove two further regimes: (II) bounded third-order skewness plus finite Fisher--Rao diameter yields a dimension-independent bound; and (III) for products of scale families---including Gaussian covariances and Gamma rates---\(R\) grows as \(Θ(\sqrt{\log d})\), unbounded in \(d\). Under a per-step Fisher-chord budget, \(R\) translates to a practical computational cost: NGD requires asymptotically at least \(R\) times as many steps as an optimizer following the Fisher--Rao geodesic. Experiments confirm all three regimes: \(R=1\) to machine precision for quadratic-potential families (I), the categorical bound \(π/(2\sqrt{2})\) is approached but not attained (II), and sampled scale-product \(R\) grows with \(d\), reaching \(R \approx 1.5\) for long, high-dimensional moves (III).
发表机构
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。