发表机构
Australian National University; CSIRO Data61; Amazon(澳大利亚国立大学; 联邦科学与工业研究组织数据61实验室; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何让预训练扩散模型的去噪动力学适应判别表示学习,提出D³CL方法,通过结合对比目标与去噪重建损失,利用LoRA更新保持轻量级,实验证明该方法能使重建与对比目标互补,是高效参数适应框架。
AI 中文摘要
文本到图像的扩散模型展现出前所未有的生成能力,且包含丰富中间表示,可用于判别视觉任务。本文研究如何在参数高效更新下,使预训练扩散模型的去噪动力学适应判别表示学习并保留生成行为,提出D³CL。关键观察是不同扩散时间步的噪声潜在表示可视为同一图像的随机视图,据此将对比目标与标准去噪重建损失结合。通过LoRA更新预训练的Stable Diffusion主干并冻结原模型参数来保持轻量级适应。实验表明重建和噪声水平对比目标可互补,D³CL是预训练扩散模型的高效参数适应框架。
英文摘要
Text-to-image diffusion models exhibit unprecedented generative capability and contain rich intermediate representations that can be useful for discriminative vision tasks. Motivated by this observation, we study a focused question: how can the denoising dynamics of a pretrained diffusion model be adapted to support discriminative representation learning while preserving its generative behavior under parameter-efficient updates? We present D$^3$CL as an investigation of this question. Our key observation is that noisy latents at different diffusion timesteps can be interpreted as stochastic views of the same underlying image, enabling a contrastive objective to be coupled with the standard denoising reconstruction loss. This formulation provides a simple way to probe the interaction between generative denoising and discriminative representation learning without training from scratch. To keep the adaptation lightweight, we apply LoRA updates to a pretrained Stable Diffusion backbone while freezing the original model parameters. D$^3$CL provides strong empirical evidence that reconstruction and noise-level contrastive objectives can be complementary: on ImageNet-1K, it obtains 80.1% linear-probing accuracy and an FID of 5.56 for $256 \times 256$ unconditional generation. Additional ablations on the design space suggest that the usefulness of diffusion features depends on where and how denoising states are sampled. These results establish D$^3$CL as a parameter-efficient adaptation framework for pretrained diffusion models, showing that noise-level contrastive learning can structure denoising representations for discriminative tasks while maintaining generative performance.