arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高维潜在变量的可扩散性

On the Diffusibility of High-Dimensional Latents

Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li

arXiv 2609.28473首次发表:更新:

发表机构

Cornell University; Adobe; Virginia Tech; University of Washington; University of Maryland(康奈尔大学; Adobe; 弗吉尼亚理工大学; 华盛顿大学; 马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现表示自编码器微调后虽重建更好但有效维度降低,标准速度预测需拟合正交噪声致效率低下,改用x0预测聚焦信号流形,在多个强重建编码器上持续提升文生图性能。

AI 中文摘要

表示自编码器(RAEs)使扩散模型能够在预训练视觉编码器的特征空间中运行。然而,许多现成的编码器并未针对忠实重建进行优化,会丢弃细粒度的视觉细节。正如预期,对这些编码器进行微调以用于图像重建可以恢复此类细节。然而,或许与直觉相反,这一过程降低了所得表示的有效维度,且改变后的几何结构对生成产生了下游影响。具体而言,我们表明,在此高维空间中使用流匹配中的标准速度预测,要求模型拟合低维信号流形之外的正交噪声方向,从而使优化效率低下。这促使我们改用干净数据参数化($\oldsymbol{x}_{0}$-预测),其将学习重点放在底层信号流形上。在多个强重建编码器的实验中,我们表明$\oldsymbol{x}_{0}$-预测持续提升了文本到图像生成的性能。

英文摘要

Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction recovers such details. However, perhaps counterintuitively, this procedure reduces the effective dimensionality of the resulting representation, and the altered geometry has downstream effects on generation. Specifically, we show that using the standard velocity prediction in flow matching in this high-dimensional space requires the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient. This motivates using the clean data parameterization ($\boldsymbol{x}_{0}$-prediction) instead, which focuses learning on the underlying signal manifold. Across experiments with multiple strong-reconstruction encoders, we show that $\boldsymbol{x}_{0}$-prediction consistently improves text-to-image generation performance.

CommentsAccepted to ECCV 2026. Project page: https://cfeng16.github.io/on_the_diffusibility/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑