arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过内部潜变量分析对扩散模型进行统一骨干细化

Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

Haksoo Lim, Myeongjin Lee, Wonjoon Chang, Jaesik Choi

arXiv 2607.09753首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology (KAIST); INEEJI(韩国科学技术院; INEEJI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究扩散模型骨干细化问题,提出DUNE框架,通过分析内部潜变量检测突变偏差,对选定条目进行骨干特定抑制,可自然扩展到基于Transformer的模型,经实验验证该方法能提高保真度并减少幻觉。

AI 中文摘要

扩散模型在各个领域取得了显著成功,其性能与参数化得分函数的去噪骨干密切相关。本文对扩散组件进行了系统的、阶段感知分析,发现深度潜变量中的早期突变与伪像密切相关。基于此,引入了DUNE(扩散统一网络细化器),这是一个无需训练的细化框架,利用基于共享EMA的准则检测深度低噪声内部潜变量中的突变偏差,并对检测器选择的条目应用特定骨干的抑制。该原理可自然扩展到基于Transformer的扩散模型。大量实验表明,DUNE提高了保真度并减少了幻觉。

英文摘要

Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that parameterize the score function. In this paper, we present a systematic, phase-aware analysis of diffusion components and show that abrupt, early-stage fluctuations in deep latents are strongly associated with artifacts. Guided by these findings, we introduce DUNE (Diffusion Unified Network refiNEr), a training-free refinement framework that detects abrupt deviations in deep low-noise internal latents using a shared EMA-based criterion, and applies backbone-specific suppression to the detector-selected entries. Although derived from U-Net, the same detect-suppress principle extends naturally to Transformer-based diffusion models by acting on the latents of deep self-attention blocks. Extensive experiments across multiple backbones indicate that DUNE improves fidelity while reducing hallucinations, offering new insight into where and when diffusion backbones should be controlled.

Comments45 pages, 23 figures. Accepted at the European Conference on Computer Vision (ECCV) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑