发表机构
Peking University; Nanjing University; Stanford University; Cornell University; University of Oulu(北京大学; 南京大学; 斯坦福大学; 康奈尔大学; 奥卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探讨像素空间扩散变换器,针对潜在扩散模型的局限,研究直接对原始像素建模的像素空间扩散方法,介绍其在高维建模中的挑战与多模态建模优势,从多方面回顾pDiTs,总结方法、识别挑战并展望未来方向。
AI 中文摘要
潜在扩散模型(LDMs)通过在VAE压缩的潜在空间中去噪来实现高效的高分辨率图像合成。然而,固定视觉tokenizer会丢弃精细纹理和结构细节,且单独的表示和扩散训练会导致重建与生成目标不匹配。这些限制使人们重新关注像素空间扩散,它直接对原始像素建模,消除VAE瓶颈并支持端到端优化。这虽更符合高保真生成需求,但在高维建模中带来挑战。像素空间建模也为统一多模态系统提供了基础。本文从模型架构、连续生成机制和统一多模态建模角度回顾像素空间扩散变换器(pDiTs),总结方法、识别挑战并探讨未来方向。
英文摘要
Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate representation and diffusion training creates a mismatch between reconstruction and generation objectives. These limitations have renewed interest in pixel-space diffusion, which models raw pixels directly, removes the VAE bottleneck, and supports end-to-end optimization. This formulation better matches the demands of high-fidelity generation but introduces challenges in high-dimensional modeling, including noise scheduling, loss weighting, token efficiency, and scalable architecture design. Pixel-space modeling also offers a promising basis for unified multimodal systems: raw pixels, text, and task conditions can be represented in a shared token space and jointly processed by a single Transformer, narrowing the gap between visual understanding and generation. This paper reviews Pixel-Space Diffusion Transformers (pDiTs) from the perspectives of model architecture, continuous generative mechanisms, and unified multimodal modeling. We summarize representative methods, identify key technical challenges, and discuss future directions toward high-fidelity, end-to-end vision foundation models that integrate generation and understanding.