arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00626cs.CV

生成难度究竟源于何处?对目标表征的实证研究

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

Marcel Plocher, Bernhard Schölkopf, Andreas Geiger, Gege Gao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过统一模型实证分析原始像素、SD-VAE隐变量等不同目标表征对生成模型的影响,发现目标表征会将生成难度重新分配到上下文建模等环节,而非由压缩性等单一因素决定。

中文摘要 AI 辅助

目标表征定义了图像生成器必须学习的分布,但常被视为可互换的接口。这一假设对连续掩码生成器而言尤为存疑,该生成器将可见token的上下文推理与每个缺失token的条件建模相结合。我们在统一的掩码自回归修正流模型中研究了原始像素、SD-VAE隐变量、DINOv2以及MAE表征自编码器特征。在共享的ImageNet训练预算下,这些空间呈现出截然不同的优化与推理模式:DINOv2在迭代次数和计算量上收敛最快,但更受益于更宽的局部去噪器与直接上下文融合;像素的优化速度显著更慢,且需要不同的预测、掩码和引导配置;MAE重构图像更忠实且呈现清晰的语义聚类,但生成效果远差于DINOv2。这些表征对无分类器引导的响应也不同,且占据着截然不同的精确率-召回率权衡。综上,我们的结果表明,压缩性、重构保真度、token维度和可见语义聚类均无法单独预测生成行为,相反,目标表征会将难度重新分配到上下文建模、逐token去噪和推理时的分布控制中。

英文摘要

The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumption is particularly questionable for continuous masked generators, which combine contextual inference from visible tokens with conditional modeling of each missing token. We study raw pixels, SD-VAE latents and DINOv2 as well as MAE representation-autoencoder features within a unified masked autoregressive rectified-flow model. Under a shared ImageNet training budget, these spaces exhibit distinct optimization and inference regimes. DINOv2 converges fastest in both iterations and computation but benefits strongly from a wider local denoiser and direct context fusion. Pixels optimize substantially more slowly and require a different prediction, masking, and guidance configuration. MAE reconstructs images more faithfully and exhibits clear semantic clustering, yet produces generations substantially worse than DINOv2. The representations also respond differently to classifier-free guidance and occupy distinct precision-recall trade-offs. Together, our results show that compression, reconstruction fidelity, token dimensionality, and visible semantic clustering do not individually predict generative behavior. Instead, target representations redistribute difficulty across contextual modeling, per-token denoising, and inference-time distributional control.

补充信息

↑