arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

像素扩散模型的对抗训练

Adversarial Training for Pixel Diffusion

Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen

arXiv 2609.38170首次发表:更新:

发表机构

UC San Diego; Adobe Research; UC Merced(加州大学圣迭戈分校; Adobe 研究院; 加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出对像素扩散模型进行对抗训练后训练,在不改变架构和采样的情况下,通过添加对抗损失恢复缺失的高频功率,联合提升分布保真度、覆盖率、提示对齐和感知质量,并指出直接输出访问是关键因素。

AI 中文摘要

像素扩散模型直接生成RGB图像,避免了自编码器的瓶颈,但其输出仍然系统性地低估了细粒度自然图像统计特征。我们证明对抗学习为这一缺陷提供了有效的训练后修正。从预训练模型出发,我们保留其原有的扩散或流匹配目标,并在非高噪声时间步对预测输出添加对抗损失,保持模型架构和采样过程不变。据我们所知,这是首个针对像素扩散模型对抗训练后训练的系统性研究。在两种像素骨干网络上,该方法联合提升了分布保真度、覆盖率、提示对齐和感知质量。我们进一步探究其工作原理。频带和幂律分析表明,原始模型系统性地低估了自然图像的高频内容,而对抗训练后训练恢复了缺失的频谱功率。相比之下,感知损失虽然也增加了高频内容,但牺牲了分布保真度和提示对齐。最近邻、召回率和匹配的无GAN SFT对照进一步排除了记忆化、模式坍缩和额外优化作为简单解释。最后,我们检验了该效应的边界。在测试的潜在扩散配置下,相同程序未产生可比的联合增益,且几乎未增加解码后的高频功率。这些结果确定了直接输出访问被修正的图像统计特征是决定对抗训练后训练成功与否的关键因素。

英文摘要

Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. Frequency-band and power-law analyses show that the original models systematically underproduce natural-image high-frequency content, while adversarial post-training restores this missing spectral power. In contrast, perceptual loss also increases high-frequency content but sacrifices distribution fidelity and prompt alignment. Nearest-neighbor, recall, and matched no-GAN SFT controls further rule out memorization, mode dropping, and additional optimization as simple explanations. Finally, we examine the boundary of this effect. Under the tested latent diffusion configurations, the same procedure does not produce comparable joint gains and adds almost no decoded high-frequency power. These results identify direct output access to the image statistics being corrected as a key factor governing when adversarial post-training succeeds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑