发表机构
Nanjing Normal University; Nanjing University of Chinese Medicine(南京师范大学; 南京中医药大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现可微学习中前向损失与反向梯度可能不一致,提出AR/FR固定深度解析实现,联合构造势函数与导数,保证前后向一致性并提供三次阶误差界。
AI 中文摘要
可微分学习通常假设前向传播中评估的标量目标与反向传播中提供给优化器的梯度描述的是同一个数学对象。我们表明,当概率目标依赖于有限特殊函数递推、自定义反向规则和数值截断时,这种对应关系可能失效。在高维von Mises-Fisher学习中,真实的数值实现可以在相同的学习状态下产生相同的前向分数和损失,同时提供不同的梯度并遵循不同的优化轨迹。我们在有限起点贝塞尔递推中刻画了这种不匹配的结构,并表明类级径向不匹配可以通过概率组合成局部非保守的更新场。因此,单独评估特殊函数值和导数的精度不足以刻画实际实现的学习目标。受此观察启发,我们引入了AR/FR,一种固定深度的解析实现,它联合构造势函数及其导数,通过构造确保前后向一致性。我们在整个非负浓度轴上建立了相对于精确贝塞尔比率的统一三次阶误差界,并将这一保证推广到学习分数和目标。随着表示维度的增加,原始有限递推变得顺序更深,而AR/FR的最坏情况误差保证以三次方式收紧,共同提供一致性、认证保真度和固定深度计算。这些结果表明,一个可微数值原语由它实现的值和它实际提供给优化器的导数共同定义;它们共同构成了学习算法的数值实现。
英文摘要
Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the backward pass describe the same mathematical object. We show that this correspondence can fail when probabilistic objectives rely on finite special-function recurrences, custom backward rules, and numerical clipping. In high-dimensional von Mises-Fisher learning, real numerical implementations can produce identical forward scores and losses at the same learning state while supplying different gradients and following different optimization trajectories. We characterize the structure of this mismatch in finite-start Bessel recurrence and show that classwise radial mismatch can compose through probabilities into a locally nonconservative update field. Evaluating the accuracy of special-function values and derivatives separately is therefore insufficient to characterize the realized learning objective. Motivated by this observation, we introduce AR/FR, a fixed-depth analytic realization that constructs a potential and its derivative jointly, ensuring forward-backward coherence by construction. We establish a uniform cubic-order error bound relative to the exact Bessel ratio over the entire nonnegative concentration axis and propagate this guarantee to learning scores and objectives. As representation dimension increases, the original finite recurrence becomes sequentially deeper, whereas the worst-case AR/FR error guarantee tightens cubically, jointly providing coherence, certified fidelity, and fixed-depth computation. These results suggest that a differentiable numerical primitive is defined by both the values it realizes and the derivatives it actually supplies to the optimizer; together, they constitute the numerical realization of the learning algorithm.
Comments30 pages