arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10524cs.CV

GRACE:面向高效视频生成的生成感知潜空间压缩

GRACE: Generation-aware latent compression for efficient video generation

Jiyoung Kim, Paul Hyunbin Cho, Jisu Nam, Donghoon Lee, Hyunsung Go, Yeonkyeong Lee, Hansaem Kim, Seungryong Kim

首次发表
浏览论文内容

中文总结 AI 辅助

GRACE通过两阶段框架压缩预训练视频自编码器,保留冻结基础潜变量并学习残差,对齐生成特征,实现8倍令牌减少和11.1倍加速,同时保持生成质量。

中文摘要 AI 辅助

高压缩率的视频自编码器提供了一种加速视频扩散模型的有效方式,因为扩散变换器(DiT)处理的是数量大幅减少的令牌。然而,此类自编码器训练难度较大,因为更高的压缩比会降低重建质量,而恢复质量需要更多的通道,这已知会减慢DiT的收敛速度。压缩后的潜变量也与DiT训练时所用的潜变量不同,因此预训练的DiT要么需要从头重新训练,要么需要以相当大的成本进行适配。压缩DiT训练时所用的自编码器似乎能保持兼容性,但仅针对重建进行优化仍会使潜变量偏离DiT所学到的分布。为解决这一问题,我们提出了面向高效视频生成的生成感知潜空间压缩(GRACE),这是一个两阶段框架,在压缩预训练视频自编码器的同时保持其与预训练DiT的兼容性。具体来说,我们保留预训练编码器中的冻结基础潜变量,并学习一个残差潜变量来补偿在更强压缩下丢失的信息,同时在冻结DiT的特征空间中将压缩后的潜变量与预训练潜变量对齐,从而使自编码器针对生成进行优化。随后,我们通过轻量级微调和非对称去噪来适配DiT,其中基础部分先于残差部分进行去噪。GRACE将Wan2.1-I2V-14B的令牌数量减少了8倍,在480x832x81分辨率下延迟降低了11.1倍,同时在VBench上保持了压缩前预训练管道的生成质量。

英文摘要

Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.

发表机构

  • KAIST AI(韩国科学技术院人工智能学院)
  • Kakao Corp.(Kakao公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑