PixelDiT2:基于表示锚定的像素扩散Transformer
PixelDiT2: Representation-Grounded Pixel Diffusion Transformers
浏览论文内容
中文总结 AI 辅助
PixelDiT2通过冻结的视觉基础模型提供显式逐块表示指导,解耦表示学习与像素生成,在ImageNet上以更少周期达到更优FID。
中文摘要 AI 辅助
近期像素空间扩散模型的进展缩小了其与潜在空间扩散在图像质量上的差距,但像素扩散模型仍收敛更慢,最终图像质量也落后。我们认为关键原因在于缺乏显式的表示先验:与通常在紧凑且结构化的潜在空间中降噪的潜在扩散不同,像素扩散需要从原始RGB空间同时学习利于降噪的表示和像素生成。为解决此问题,我们提出PixelDiT2,一种端到端的像素空间扩散模型,旨在将表示学习与像素生成解耦,且不引入自编码器或潜在重建瓶颈。我们提出表示锚定方法,使用冻结的预训练视觉基础模型在整个降噪过程中提供显式的逐块表示指导,使像素扩散Transformer能更专注于像素生成。在ImageNet-256x256上,PixelDiT2在600个训练周期后达到FID 1.46;在512x512分辨率下,PixelDiT2在680个训练周期后达到FID 1.48。
英文摘要
Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which usually denoises in a compact and structured latent space, pixel diffusion needs to learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose representation grounding that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-256x256, PixelDiT2 achieves an FID of 1.46 after 600 epochs; at 512x512 resolution, PixelDiT2 achieves an FID of 1.48 after 680 epochs. Project page: https://pixeldit.github.io/pixeldit2/
发表机构
- NVIDIA(英伟达)
- University of Rochester(罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。