牛顿匹配用于生成建模:微调与采样的统一框架
Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling
浏览论文内容
中文总结 AI 辅助
提出牛顿匹配统一框架,通过迭代优化规范模型和 Fisher-Rao 度量实现生成模型的微调与采样,证明反向 KL 下降与二次收敛,并统一多种现有方法。
中文摘要 AI 辅助
我们开发了牛顿匹配(Newton Matching),这是一个用于生成建模中微调和采样的统一框架。目标是 $\pi\propto\mu e^{\tau r}$,其中 $r$ 是奖励,$\tau>0$ 是逆温度,$\mu$ 表示预训练模型的终端密度(用于微调)或常数 $1$(用于采样)。我们将范式从孤立损失转移到对规范模型的迭代优化:标准条件匹配的总体最小化器用于终端密度。在兼容的光滑实现假设下,规范速度构成一个与密度流形微分同胚的流形。将 Fisher-Rao 度量与混合连接传输到该流形,我们证明反向 KL 海森矩阵等于该度量,因此牛顿方向与负 Fisher-Rao 梯度一致。在终端密度 $\rho$ 处,每个阶段采取由正则化奖励 $r-\frac1\tau\log(\rho/\mu)$ 生成的切向步长,随后进行保持终端密度的规范化。这种规范回缩产生了精确的有限步长密度表征。对于理想迭代,我们证明了在 $0 < \eta \le \tau$ 时严格反向 KL 下降(远离目标),在温和条件下的全局收敛性,以及全步长($\eta=\tau$)的局部二次收敛。协方差和梯度形式,每种都有前向或反向回归对构造,产生具有相同总体最小化器的样本级切向更新损失,无需重要性采样或全轨迹反向传播。我们开发了近似更新,并将临界点一致性定义为切向位移消失当且仅当 $\rho=\pi$。我们将代表性方法恢复为精确实现、临界点一致近似或目标改变变体,从而实现模块化算法设计。我们的工作推进了生成模型强化学习的理论和算法。
英文摘要
We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. The target is $π\proptoμe^{τr}$, where $r$ is the reward, $τ>0$ the inverse temperature, and $μ$ denotes the pretrained model's terminal density for fine-tuning or the constant $1$ for sampling. We shift the paradigm from isolated losses to iterative optimization over canonical models: population minimizers of standard conditional matching for terminal densities. Under compatible smooth-realization assumptions, canonical velocities form a manifold diffeomorphic to the density manifold. Transporting the Fisher-Rao metric and mixture connection to this manifold, we show that the reverse-KL Hessian equals the metric, so the Newton direction coincides with the negative Fisher-Rao gradient. At terminal density $ρ$, each stage takes a tangential step generated by the regularized reward $r-\frac1τ\log(ρ/μ)$, followed by terminal-density-preserving canonicalization. This canonical retraction yields an exact finite-stepsize density characterization. For the ideal iteration, we prove strict reverse-KL descent away from the target for $0 < η\le τ$, global convergence under mild conditions, and local quadratic convergence for full steps ($η=τ$). Covariance and gradient forms, each with forward or reverse regression-pair constructions, yield sample-wise tangential-update losses with the same population minimizer, without importance sampling or full-trajectory backpropagation. We develop approximate updates and define critical-point consistency as vanishing tangential displacement if and only if $ρ=π$. We recover representative methods as exact realizations, critical-point-consistent approximations, or objective-altering variants, enabling modular algorithm design. Our work advances the theory and algorithms of reinforcement learning for generative models.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。