arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38156cs.CV

DMA$^2$:基于对抗与锚点损失的像素空间分布匹配

DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses

  • UC San Diego(加州大学圣迭戈分校)
  • Adobe Research(Adobe研究院)
  • University of Virginia(弗吉尼亚大学)
  • UC Merced(加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Shaoteng Liu, Lehan Yang, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen

AI总结:

提出DMA$^2$,通过固定高噪声匹配带、DINO-Adv对抗指导和AF-Loss分布目标,实现像素空间少步扩散蒸馏,四步生成超越25步教师。

AI中文摘要:

分布匹配蒸馏(DMD)为少步扩散生成提供了通用框架,但其现代文本到图像实例主要围绕潜在扩散开发,因此忽略了原生RGB的关键特性和设计机会。我们重新审视了像素空间教师的两个DMD接口。在教师匹配方面,诊断显示低噪声RGB匹配主要由局部纹理线索主导,这促使我们采用固定的高噪声匹配带。在真实数据方面,原生干净RGB输出允许从外部视觉表示获得指导,而无需穿越解码器或共享重型伪评分判别器。DINO-Adv从对抗梯度路径中移除此判别器,并提供局部参数化补丁指导。对于分布级指导,我们引入了AF-Loss,一种为文本到图像DMD设计的无参数辅助语义分布场目标。它在共享的DINOv2空间中作用于分离的滚动真实和生成支持,同时保留提示条件的教师监督。AF-Loss不增加可学习参数或推理时计算。这些设计共同构成DMA$^2$。在DPG-Bench、GenEval、VQAScore和COCO30K上,四步DMA$^2$学生模型优于25步教师模型和评估的少步蒸馏器。

英文摘要:

Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA$^2$. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA$^2$ student performs better than the 25-step teacher and evaluated few-step distillers.

↑