arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于观测算子的像素空间扩散模型

Pixel-Space Diffusion via Observation Operators

Shaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang

arXiv 2608.21885首次发表:更新:

AI 中文总结

针对像素空间扩散模型的尺度-时间不匹配问题,提出观测算子扩散框架,采用时变观测轨迹与GL-CoDA解码器,在ImageNet-256上FID达1.52,收敛更快且生成质量提升。

AI 中文摘要

像素空间扩散模型直接对图像分布进行建模,但优化难度较大。现有方法通过目标重参数化缓解该挑战,但在去噪过程中仍依赖固定的干净图像目标。通过实证分析,我们发现了尺度-时间不匹配问题:随着噪声降低,图像结构从粗到细变得可预测,而现有模型即便在高噪声下也需预测完整图像,导致低信噪比梯度阻碍优化。为解决该不匹配问题,我们提出Observation Operator Diffusion(观测算子扩散),这是一个统一框架,将监督轨迹和特征细化与图像结构的固有恢复顺序对齐。具体而言,我们将标准流路径上的固定全图监督替换为时间索引的观测轨迹,该轨迹在去噪过程中从粗结构演变为全图。该轨迹由一系列不同观测尺度的Gaussian-Lanczos算子实例化,生成路径一致的训练目标。我们进一步引入GL-CoDA,这是一个解码器,在解码阶段注入特定尺度的Gaussian-Lanczos观测以实现从粗到细的特征细化。大量实验表明,所提方法收敛速度显著加快,同时持续提升生成质量,在ImageNet-256上取得了1.52的FID。

英文摘要

Pixel-space diffusion models directly model image distributions but remain difficult to optimize. Recent methods alleviate this challenge through target reparameterization, while still relying on a fixed clean-image target throughout denoising. Through empirical analysis, we identify a scale-time mismatch: image structures become predictable from coarse to fine as noise decreases, whereas existing models are forced to predict the full image even under high noise, resulting in low-SNR gradients that hinder optimization. To resolve this mismatch, we propose Observation Operator Diffusion, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures. Specifically, we replace fixed full-image supervision along the standard flow path with a time-indexed observation trajectory that evolves from coarse structures to the full image during denoising. This trajectory is instantiated with a family of Gaussian-Lanczos operators at varying observation scales, yielding a path-consistent training objective. We further introduce GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine feature refinement. Extensive experiments show that the proposed approach converges substantially faster while consistently improving generation quality, achieving an FID of 1.52 on ImageNet-256.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑