arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39222cs.CV

DC-SAE:用于加速扩散模型收敛的深度压缩语义自编码器

DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence

Xu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

DC-SAE通过解耦语义编码器与像素级编码器,实现高压缩比的同时加速扩散模型训练收敛,在ImageNet上取得32倍压缩和更优的重建质量与生成性能。

中文摘要 AI 辅助

高压缩分词器对于扩展潜在图像生成模型至关重要。然而,激进的压缩在重建保真度和生成效率之间造成了根本性的权衡:高压缩图像编码器总是增加扩散训练的学习难度,导致模型收敛缓慢。最近的表示自编码器通过用预训练语义编码器替换VAE编码器来提升潜在特征表达能力,从而加速扩散训练,但它们通常仅限于中等压缩,并且丢失了忠实重建所需的像素级细节。为了同时实现高压缩和快速扩散训练,我们提出了DC-SAE,一种专为高压缩图像生成和加速扩散模型收敛而设计的解耦紧凑语义自编码器。DC-SAE由两个关键组件组成:(1)利用语义编码器实现更高压缩比的宏观架构设计,以及(2)保留低级细节的像素级编码器,确保高保真图像重建。我们通过实验证明,DC-SAE在图像生成任务上表现强劲,同时实现了紧凑的潜在表示和高效的训练动态。具体来说,在ImageNet数据集上,分辨率为$512 \ imes 512$时,DC-SAE实现了$32\ imes$的空间压缩,PSNR为29.79,gFID为3.37,在PSNR和gFID上分别大幅优于之前最先进的高压缩分词器基线DC-AE达13.5%和54.9%,同时保持相当的吞吐量和更快的扩散模型训练收敛速度。除了类条件生成之外,使用DC-SAE的$1.6$B参数DiT在$1024\ imes1024$分辨率下进行文本到图像生成,在GenEval上达到0.84,在DPG-Bench上达到86.007。

英文摘要

High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion training, resulting in slow model convergence. Recent representation autoencoders speed up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. Specifically, on the ImageNet dataset with $512 \times 512$ resolution, DC-SAE achieves $32\times$ spatial compression, with 29.79 PSNR and 3.37 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively, maintaining comparable throughput and faster diffusion model training convergence. Beyond class-conditional generation, a $1.6$B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at $1024\times1024$ resolution.

发表机构

  • Peking University(北京大学)
  • Singapore University of Technology and Design(新加坡科技设计大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑