arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.12013cs.CVcs.AI

L2P:解锁像素生成的潜在能力

L2P: Unlocking Latent Potential for Pixel Generation

  • Nanjing University(南京大学)
  • Tencent Youtu Lab(腾讯云图实验室)
  • Hainan-biuh(海南-比乌)
  • Weess Gmbh(韦斯公司)

机构由 AI 辅助整理,请以论文原文为准。

Zhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang, Jiawei Chen, Zhuoqi Zeng, Wei Zhang, Chengjie Wang, Jian Yang, Ying Tai

更新

AI总结:

本文提出L2P框架,通过利用预训练LDM的知识,高效构建像素空间模型,无需大量数据和计算资源,实现快速收敛和高分辨率生成。

AI中文摘要:

像素扩散模型近期重新受到视觉生成的关注。然而,从头训练高级像素空间模型需要极高的计算和数据资源。为此,我们提出Latent-to-Pixel(L2P)转移范式,一种高效的框架,直接利用预训练LDM的丰富知识构建强大像素空间模型。具体而言,L2P摒弃VAE,采用大块标记化并冻结源LDM的中间层,仅训练浅层以学习潜在到像素的转换。通过使用LDM生成的合成图像作为唯一训练数据集,L2P拟合已平滑的数据流,实现无需真实数据收集的快速收敛。该策略使L2P能够仅用8块GPU无缝迁移大量潜在先验到像素空间。此外,消除VAE内存瓶颈解锁了原生4K超高清分辨率生成。在主流LDM架构上的广泛实验表明,L2P的训练开销极小,但在DPG-Bench上与源LDM表现相当,在GenEval上达到93%的性能。

英文摘要:

Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohibitive computational and data resources. To address this, we propose the Latent-to-Pixel (L2P) transfer paradigm, an efficient framework that directly harnesses the rich knowledge of pre-trained LDMs to build powerful pixel-space models. Specifically, L2P discards the VAE in favor of large-patch tokenization and freezes the source LDM's intermediate layers, exclusively training shallow layers to learn the latent-to-pixel transformation. By utilizing LDM-generated synthetic images as the sole training corpus, L2P fits an already smooth data manifold, enabling rapid convergence with zero real-data collection. This strategy allows L2P to seamlessly migrate massive latent priors to the pixel space using only 8 GPUs. Furthermore, eliminating the VAE memory bottleneck unlocks native 4K ultra-high resolution generation. Extensive experiments across mainstream LDM architectures show that L2P incurs negligible training overhead, yet performs on par with the source LDM on DPG-Bench and reaches 93% performance on GenEval.

补充信息

↑