发表机构
LLM Suite Team, JP Morgan Chase & Co.(摩根大通 LLM Suite 团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究生成流训练损失的意义,证明差值型损失可界定采样误差,比值型损失不可;并揭示反向策略决定收敛性,梯度下降沿其扩散,实现全局收敛。
AI 中文摘要
生成流通过训练一个流使其达到平衡,从未归一化目标中采样,训练损失是实践者关注的信号。我们探究该信号的价值:小的损失是否能保证采样器的准确性,损失能否被驱动至零,以及梯度下降实现这一目标的速度。损失决定了第一个问题。通过差值比较平衡两侧的损失,以全变差界定了流所隐含的采样器的误差,其显式常数不涉及策略;而通过比值比较的流匹配损失,在单个环上,只要其生成器在平衡处连续,就不存在这样的界。在图上,反向策略决定了另外两个问题。一旦反向策略被冻结,平衡就变为反向链下的不变性,因此在有限图上存在性是自由的,一个常数——反向链格林算子的范数,扮演逆谱间隙的角色——从上下两个方向确定了平衡流附近损失曲率的阶数,并为训练在其附近的收敛速度设定了下限。其机制是梯度下降沿反向策略扩散流。对于详细平衡和轨迹平衡的平方对数生成器,在状态上训练平衡损失,在每个有限路径连通图上,从任意正初始化出发,都能全局收敛。该常数可以是无穷大,而反向轨迹平均较短,此时精确流匹配可能失败。这些界和速率通过在可枚举状态空间上的精确计算得到验证,每个定理都带有从Lean 4开发中计算出的认证状态。
英文摘要
Generative flows sample from an unnormalized target by training a flow to be balanced, and the training loss is the signal a practitioner watches. We ask what that signal is worth: whether a small loss certifies an accurate sampler, whether the loss can be driven to zero, and how fast gradient descent does so. The loss decides the first. Losses that compare the two sides of the balance by their difference bound, in total variation, the error of the sampler the flow implies, with explicit constants that do not involve the policy; flow-matching losses that compare them through a ratio admit no such bound, already on a single cycle, whenever their generator is continuous at balance. On graphs, the backward policy decides the other two. Once it is frozen, balance becomes invariance under the backward chain, so that existence is free on finite graphs, and one constant --- the norm of that chain's Green operator, which plays the role of an inverse spectral gap --- fixes the order of the curvature of the loss around the balanced flow, from above and below, and sets a floor under the rate at which training converges near it. The mechanism is that gradient descent diffuses the flow along the backward policy. For the squared-logarithm generator of detailed and trajectory balance, training the balance loss on states converges globally on every finite path-connected graph, from every positive initialization. The constant can be infinite while backward trajectories are short on average, and exact flow matching can then fail. The bounds and rates are tested by exact computation on enumerable state spaces, and every theorem carries a certification status computed from a Lean~4 development.
Commentspending corporate approval