arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冻结的像素空间扩散模型可利用自身样本进行自引导

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

Zixuan Fu, Chong Wang, Lanqing Guo, Kailai Zhou, Jiahao Nie, Bihan Wen

arXiv 2607.29122首次发表:更新:

AI 中文总结

本研究提出合成自引导(SSG)策略,通过冻结预训练像素扩散模型主干、附加轻量预测头,利用模型自身样本训练头,以预测差异为自引导方向,大幅提升像素扩散模型生成效果且训练成本极低。

AI 中文摘要

像素空间扩散模型旨在直接在原始像素上学习端到端生成器,这颇具挑战性,因为单个模型需在同一高维空间中同时捕捉全局结构与局部纹理。尽管近期研究通过采用不同预测目标、训练目标及架构改进了像素扩散,但这些进展通常需要从头训练新模型。本文提出一种更经济且互补的策略:**冻结的预训练像素扩散模型可实现自引导**。核心观察是,预训练像素扩散Transformer的中间层可解码为捕捉主要低频结构的粗预测,而最终层则逐步细化局部高频细节。因此,我们在中间层附加一个轻量预测头,保持主干冻结,并在采样过程中利用中间层与最终层预测的差异作为自引导方向。训练该头时,我们进一步发现无需真实图像,模型生成的样本已足够,且在训练该头时甚至优于真实图像,尤其能增强像素扩散易欠拟合的高频分量。在ImageNet上的多个像素扩散模型中,我们的**合成自引导(Synthetic Self-Guidance, SSG)**持续提升生成效果,且适配器训练所需计算量不足完整模型训练的1%:在无分类器引导(Classifier-Free Guidance, CFG)的评估JiT变体中,FID降低超50%;在带CFG的强基准模型中进一步提升,例如JiT-H/16从1.86降至1.67,PixelREPA-H/16从1.81降至1.59。代码可在该https URL获取。

英文摘要

Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both global structure and local texture in the same high-dimensional space. While recent work improves pixel diffusion through alternative prediction targets, training objectives, and architectures, these advances typically require training a new model from scratch. We show there is a cheaper, complementary strategy: \textbf{a frozen, pretrained pixel diffusion model can guide itself}. Our key observation is that intermediate layers of a pretrained pixel diffusion transformer can be decoded into coarse predictions that capture the main low-frequency structure, while the final layers progressively refine local, high-frequency details. We therefore attach a lightweight prediction head to an intermediate layer, keep the backbone frozen, and use the discrepancy between the intermediate and final predictions as a self-guidance direction during sampling. To train this head, we further find that real images are not necessary. Instead, model-generated samples suffice and even outperform real images for training the head, especially in enhancing the high-frequency components that pixel diffusion tends to underfit. Across multiple pixel diffusion models on ImageNet, our \textbf{Synthetic Self-Guidance (SSG)} consistently improves generation while adapter training requires less than 1$\%$ of full-model training compute: it reduces FID by over 50$\%$ across the evaluated JiT variants without classifier-free guidance (CFG) and further improves strong baselines with CFG, e.g., JiT-H/16 from 1.86 to 1.67 and PixelREPA-H/16 from 1.81 to 1.59. Our code is available at https://github.com/zfu006/SSG.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑