arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33499cs.LGcs.ITmath.ITmath.OC

有用自然梯度更新的代价

The cost of useful natural gradient updates

Subhransu S. Bhattacharjee, Dylan Campbell, Rahul Shome

AI总结:

本文研究自然梯度更新中步长而非方向决定计算代价,证明在KL预算下有用步骤的样本复杂度及NP难性,并给出实验验证。

AI中文摘要:

将自然梯度方向转化为有用的有限更新需要哪些信息?在总体库尔贝克-莱布勒(KL)预算下,如果一步可行且损失不超过沿该方向最佳可行增益的ε比例,我们称之为有用步骤。我们构造了一个四状态指数族,其各分布的初始梯度、标量Fisher信息量和自然梯度相同,但其中两个分布的有用步骤集合不相交。当这些量被精确提供且分布仅通过抽样已知时,该族的最坏情况样本复杂度为Θ(log(1/δ)/(pε²))(ε较小),其中p缩放稀有状态概率,δ为失败概率。预算固定且最优增益保持有界远离零,因此步长而非方向承担此代价。对于简洁描述的事件倾斜模型,即使具有精确自然梯度和高效精确采样,返回有用步骤也是NP难的。即使在Fisher条件数至多为3的两参数逻辑斯蒂族中,将单位自然梯度恢复到常数误差也是NP难的。我们还给出了事件倾斜的匹配样本界、阻尼Fisher求解的样本界以及仿射分类器的总体KL证书。在冻结特征分类器头部,在抽样KL边界处停止在约一半试验中成功,而10%的KL余量在KL预算为0.01时将联合成功率提升至93%以上。因此,知道向何处移动是不够的:移动多远可能承担更新的全部代价。

英文摘要:

What information is needed to turn a natural-gradient direction into a useful finite update? Under a population Kullback-Leibler (KL) budget, we call a step useful if it is feasible and loses at most a fraction $\varepsilon$ of the best feasible gain along the direction. We construct a four-state exponential family whose laws share their initial gradient, scalar Fisher information and natural gradient, yet two laws have disjoint useful-step sets. With these quantities supplied exactly and the law otherwise known only through draws, the family's worst-case sample complexity is $Θ(\log(1/δ)/(p\varepsilon^2))$ for small $\varepsilon$, where $p$ scales rare-state probabilities and $δ$ is the failure probability. The budget is fixed and the optimal gain stays bounded away from zero, so the step length, not the direction, carries this cost. For succinctly described event-tilt models, returning a useful step is NP-hard even with the exact natural gradient and efficient exact sampling. Recovering the unit natural gradient to constant error is also NP-hard even in a two-parameter logistic family with Fisher condition number at most 3. We also give matching sample bounds for event tilts, sample bounds for damped Fisher solves and a population-KL certificate for affine classifiers. In frozen-feature classifier heads, stopping at a sampled KL boundary succeeds in about half of the trials, and a 10% KL margin raises joint success above 93% at a KL budget of 0.01. Thus, knowing where to move is not enough: how far to move can carry an update's entire cost.

补充信息

↑