带引导式稀疏全局细化的高级像素扩散模型
Advanced Pixel Diffusion Model with Guided Sparse Global Refinement
浏览论文内容
中文总结 AI 辅助
针对现有像素扩散模型的缺陷,提出 PixSGR 框架,通过引导式稀疏全局细化实现高效高质量图像生成,在 ImageNet 上 256×256 分辨率 FID 达 1.51,512×512 分辨率达 1.60。
中文摘要 AI 辅助
像素空间扩散近年来成为高保真图像生成的有前景方向,其直接在原始像素域对图像建模。然而,自然图像的极高维度使像素空间扩散计算量巨大。现有像素扩散模型为提升效率,要么用大补丁分词法牺牲细节,要么将后续细化限制在单个补丁内。这种补丁内细化必然限制补丁边界的结构连续性与长程 token 交互,降低细化质量。为解决这些问题,我们提出 PixSGR,一种专为直接在像素空间建模自然图像分布设计的新型像素扩散框架。PixSGR 从有监督低通道瓶颈开始,高效捕获自然图像的低维流形;随后逐步扩展通道维度与空间分辨率,恢复越来越精细的结构。在空间细化阶段,粗尺度注意力图预选全局相关交互,预稀疏化细尺度注意力,实现超越孤立补丁的非局部细化,且无密集注意力的二次成本。在 ImageNet 上的大量实验验证了 PixSGR 的有效性:其在 256×256 分辨率下实现 FID 1.51,且扩展至 512×512 分辨率时仍保持性能,达到 FID 1.60。
英文摘要
Pixel-space diffusion has recently emerged as a promising direction for high-fidelity image generation by modeling images directly in the original pixel domain. However, pixel-space diffusion is computationally demanding due to the extremely high dimensionality of natural images. For efficiency, existing pixel diffusion models either compromise fine details with large-patch tokenization or confine subsequent refinement within individual patches. Such intra-patch refinement inevitably restricts structural continuity across patch boundaries and long-range token interactions, limiting refinement quality. To address these issues, we propose PixSGR, a novel Pixel diffusion framework with Sparse Global Refinement tailored for modeling the distribution of natural images directly in pixel space. PixSGR starts from a supervised low-channel bottleneck to efficiently capture the low-dimensional manifold of natural images. It then progressively expands the channel dimensionality and spatial resolution to recover increasingly fine-grained structures. At the spatial refinement stage, coarse-scale attention maps preselect globally relevant interactions to pre-sparsify fine-scale attention, enabling non-local refinement beyond isolated patches without the quadratic cost of dense attention. Extensive experiments on ImageNet validate the effectiveness of PixSGR. It achieves an FID of 1.51 at 256$\times$256 and maintains performance when scaled to 512$\times$512, attaining an FID of 1.60.
发表机构
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。