一步逆过程是凸规划:扩散逆过程的贝叶斯极限校准
One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion
AI总结:
该研究提出扩散逆过程的一步隐式DDIM步骤对应贝叶斯极限下的凸规划,分析了其解的唯一性、求解器稳定性及几何收敛域,发现训练扩散模型的分数函数违反了相关协方差上限,揭示了模型的固有局限性。
AI中文摘要:
一次隐式DDIM逆过程步骤是衡量预训练扩散模型是否编码局部流形几何的最廉价探测手段,它是显式势函数的平稳条件:$x-G(x)=\nabla\Psi_t(x)$,在贝叶斯极限下该势函数是强凸的,凸性模量恰好为$e^{-h_t}$,其中$h_t$是该步骤的对数信噪比间隙——对于任意数据分布、调度方案和数据点均成立,无需流形、可达性或单峰性假设。需区分三个结论:(i)贝叶斯极限下解唯一;若存在第二个解,则说明训练得到的分数函数违反了后验协方差界$1/(1-e^{-h_t})$,这是模型误差的无假设证明;该界还使收缩率成为调度常数$\rho_g^{\star}=1-e^{-h_t}$,在标准DDPM调度中该值始终小于0.326。(ii)求解器仍可能失效:皮卡迭代是对$\Psi_t$的单位步长梯度下降,当$\lambda_{\max}(\nabla^2\Psi_t)>2$时不稳定,因此振荡不代表任何结论;阻尼到$2/\lambda_{\max}$以下可解决该问题。(iii)几何存在于收敛域中:在无量纲深度$w=r\kappa_{\max}$尺度下,振荡壳位于$w=\frac{1}{2}$(与调度无关),发散壳位于$w=1/(1+\rho_g^{\star})$,在$\\|\mathrm{II}\\|^2$中测得有限噪声修正。精确分数在三类数据上可复现这两个壳,误差在0.54%以内;我们探测的所有训练分数均未显示出壳——这是推导得出的限制而非空结果:费米窗口与模型自身训练支持的冲突倍数为3.6-$5.6\times$,训练得到的海森-利普希茨常数是数据分布曲率的2-$12\\%$,在ReLU网络中为0。最后,仅从$\mathrm{Cov}(x_0\mid x_t)\succeq0$可得到无条件上限$\sigma_t\lambda_{\max}(\mathrm{sym}\\,J)\le1$,精确分数满足该上限的误差为$3\times10^{-7}$,但在所有DDPM CIFAR-10/CelebA-HQ-256设置中均被违反,违反倍数为1.26-$4.66\times$。
英文摘要:
One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local manifold geometry. It is the stationarity condition of an explicit potential, $x-G(x)=\nablaΨ_t(x)$, strongly convex at the Bayes limit with modulus exactly $e^{-h_t}$ for the step's log-SNR gap $h_t$ $-$ for every data law, schedule and point, with no manifold, reach or unimodality hypothesis. Three consequences must be kept apart. (i) The solution is unique at the Bayes limit; a second one requires the trained score to violate the posterior-covariance bound by $1/(1-e^{-h_t})$, a hypothesis-free certificate of model error; the same bound makes contraction a schedule constant, $ρ_g^{\star}=1-e^{-h_t}<0.326$ throughout the standard DDPM schedule. (ii) The solver can still fail: Picard iteration is unit-step gradient descent on $Ψ_t$, unstable wherever $λ_{\max}(\nabla^2Ψ_t)>2$, so oscillation certifies nothing; damping below $2/λ_{\max}$ cures it. (iii) The geometry lives in the convergence domain: on the scale-free depth $w=rκ_{\max}$ the oscillation shell sits at $w=\tfrac12$, schedule-free, and the divergence shell at $w=1/(1+ρ_g^{\star})$, with a measured finite-noise correction in $\|\mathrm{II}\|^2$. Exact scores reproduce both to within $0.54\%$ on three classes; no trained score we probe shows a shell $-$ a derived limitation, not a null result: the Fermi window conflicts with the model's own training support by $3.6$-$5.6\times$, and the trained Hessian-Lipschitz constant is $2$-$12\%$ of the curvature the law reads, $0$ on a ReLU net. Finally the unconditional ceiling $σ_tλ_{\max}(\mathrm{sym}\,J)\le1$, from $\mathrm{Cov}(x_0\mid x_t)\succeq0$ alone, holds for the exact score to $3\times10^{-7}$ but is violated in all DDPM CIFAR-10/CelebA-HQ-256 settings, by $1.26$-$4.66\times$.