x预测就足够了:通过端点可解码性实现无训练加速生成
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
- School of Physical Science and Technology, Beijing University of Posts and Telecommunications(北京邮电大学物理科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究基于x预测,通过端点可解码性提出截断跳跃采样方法,无需重新训练等操作,在多个模型和基准测试中减少神经函数评估次数20 - 70%,且质量相近,实现无轨迹重新设计的推理加速。
AI中文摘要:
扩散和流匹配模型能生成高质量样本,但其常微分方程采样器通常需要数十到数百次神经函数评估(NFEs)。对于已发布的检查点,这仍是实际挑战,因为许多加速器需要通过重新训练、蒸馏或轨迹重新设计等额外设计选择和训练成本。我们基于x预测研究了一条不同路径。采样期间,标准仿射概率路径已暴露x₀信息,我们将此属性形式化为端点可解码性,并表明解码器是通常ℓ₂目标下的最小均方误差估计器E[x₀|xₜ]。这产生了截断跳跃采样(TJS),它无需重新训练、蒸馏或架构更改。在多个模型和基准测试中,它将NFEs减少20 - 70%且质量相近,还说明了端点预测无需拉直轨迹就能工作的原因,实现了无轨迹重新设计的推理加速。
英文摘要:
Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (NFEs). This remains a practical challenge for released checkpoints, since many accelerators require additional design choices and training cost through retraining, distillation, or trajectory redesign. We investigate a different route based on $x$-prediction. During sampling, standard affine probability paths already expose $x_0$ information: an intermediate state and its path velocity determine a principled estimate of the clean sample. We formalize this property as \textbf{endpoint decodability} and show that the decoder is the minimum-MSE estimator $\mathbb{E}[x_0\mid x_t]$ under the usual $\ell_2$ objective. This yields \textbf{Truncated Jump Sampling} (TJS): stop the ODE at an early-exit time $t^*$ and return the decoded $x_0$. TJS requires no retraining, distillation, or architecture change. Across SDXL, SD3.5M, Z-Image-Turbo, and three class-conditional benchmarks, it reduces NFEs by 20--70\% with near-matched quality. The analysis also shows why endpoint prediction can work without straightening the trajectory, providing inference acceleration without trajectory redesign.