仅对齐 $\m{x}$:对齐预测而非表示
Just Align $\bm{x}$: Aligning Predictions, Not Representations
浏览论文内容
中文总结 AI 辅助
针对表示对齐在像素空间扩散模型中效果不佳的问题,提出JAx方法,通过跨噪声水平对齐干净图像预测来改进训练目标,在ImageNet 256x256上持续提升FID并加速收敛。
中文摘要 AI 辅助
表示对齐已成为加速扩散训练的有效方式,但其益处并不能可靠地迁移到像素空间的干净图像预测中。在JiT中,我们发现辅助特征对齐虽能改善对语义特征的访问,却减少了对干净图像预测所需的图像变化性,从而在辅助目标与去噪任务之间造成不匹配。这表明一个不同的原则:辅助监督应改进预测目标本身,而非强加一个独立的表示目标。我们提出JAx(仅对齐x),一种跨噪声水平对齐干净图像预测的预测监督方法。JAx通过保持原始JiT输入分布的马尔可夫退化过程,将噪声更大的学生观测与更干净的观测耦合。在此耦合下,来自更干净状态的预言机预测与最优JiT目标具有相同的条件均值,而其条件目标协方差不大于后者。因此,预言机预测对齐在保持总体JiT目标(至多相差一个常数)的同时,提供了方差更低的训练目标。为使该构造在不完美的EMA教师下切实可行,JAx将真实标签监督与基于可靠性门控的耦合带相结合,该耦合带根据预测风险选择邻近的教师状态。在ImageNet 256x256上,JAx在JiT-B/16、L/16和H/16上均持续改善FID并加速收敛,且无需外部编码器或改变架构与采样过程。梯度诊断进一步显示小批量梯度方差降低,而消融实验表明这些增益不能仅由时间重加权解释。这些结果表明,预测空间监督为像素空间生成模型提供了一种简单且原理性的表示对齐替代方案。
英文摘要
Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-image prediction. In JiT, we find that auxiliary feature alignment can improve access to semantic features while reducing access to image variation needed for clean-image prediction, creating a mismatch between the auxiliary objective and the denoising task. This suggests a different principle: auxiliary supervision should improve the prediction target itself rather than impose a separate representation target. We introduce JAx (Just Align x), a prediction-supervision method that aligns clean-image predictions across noise levels. JAx couples a noisier student observation with a cleaner observation through a Markov degradation that preserves the original JiT input distribution. Under this coupling, the oracle prediction from the cleaner state has the same conditional mean as the optimal JiT target, while its conditional target covariance is no greater. Thus, oracle prediction alignment preserves the population JiT objective up to a constant while providing a lower-variance training target. To make this construction practical with an imperfect EMA teacher, JAx combines ground-truth supervision with a reliability-gated coupling band that selects nearby teacher states based on prediction risk. On ImageNet 256x256, JAx consistently improves FID and accelerates convergence across JiT-B/16, L/16, and H/16, without an external encoder or changes to the architecture or sampling procedure. Gradient diagnostics further show reduced minibatch gradient variance, while ablations demonstrate that the gains cannot be explained by time reweighting alone. These results show that prediction-space supervision provides a simple and principled alternative to representation alignment for pixel-space generative models.
发表机构
- Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。