AI 中文总结
针对图像超分辨率,提出DiMOO-SR框架。训练时用逆频率采样处理稀有令牌,推理时用空间一致性排名提高结构连贯性。实验表明该框架仅需少量并行解码步骤就能实现有竞争力的感知质量,凸显离散扩散在图像超分辨率中的潜力。
AI 中文摘要
连续扩散模型已成为逼真图像超分辨率(SR)的主导范式,但它们通常将重建表述为连续信号级去噪,并通过外部条件模块纳入语义先验。这使得利用现代多模态模型基于统一令牌的缩放范式变得不那么直接。自回归模型通过将图像建模为离散视觉令牌提供了更原生的语义表示,但其因果解码对于高分辨率重建效率低下。离散扩散通过对视觉令牌进行非因果、并行预测提供了一个有前景的中间地带。然而,由于两个特定于任务的挑战,直接将离散扩散应用于SR仍然很棘手:(1)视觉令牌的长尾分布,它不能充分代表罕见但在感知上至关重要的纹理;(2)空间不一致的并行解码,这可能会引入孤立的伪像。为了解决这些问题,我们提出了DiMOO-SR,一种用于逼真图像SR的稀有感知多模态离散扩散框架。在训练期间,逆频率采样(IFS)优先处理代表性不足但信息丰富的令牌。在推理期间,空间一致性排名(SCR)使用局部邻域一致性来细化令牌置信度,以提高结构连贯性。在广泛使用的真实世界SR基准上进行的大量实验表明,DiMOO-SR仅通过几个并行解码步骤就实现了有竞争力的感知质量,突出了离散扩散在生成图像超分辨率方面的潜力。代码将在发表时发布。
英文摘要
Continuous diffusion models have become the dominant paradigm for photo-realistic image Super-Resolution (SR), but they typically formulate reconstruction as continuous signal-level denoising and incorporate semantic priors through external conditioning modules. This makes it less direct to exploit the unified token-based scaling paradigm of modern multimodal models. Autoregressive models provide a more native semantic representation by modeling images as discrete visual tokens, yet their causal decoding is inefficient for high-resolution reconstruction. Discrete diffusion offers a promising middle ground by enabling non-causal, parallel prediction over visual tokens. However, directly adapting discrete diffusion to SR remains non-trivial due to two task-specific challenges: (1) the long-tailed distribution of visual tokens, which under-represents rare but perceptually critical textures; and (2) spatially inconsistent parallel decoding, which may introduce isolated artifacts. To address these issues, we propose DiMOO-SR, a rarity-aware multimodal discrete diffusion framework for photo-realistic image SR. During training, Inverse Frequency Sampling (IFS) prioritizes under-represented but information-rich tokens. During inference, Spatial Consistency Ranking (SCR) refines token confidence using local neighborhood agreement to improve structural coherence. Extensive experiments on widely used real-world SR benchmarks demonstrate that DiMOO-SR achieves competitive perceptual quality with only a few parallel decoding steps, highlighting the potential of discrete diffusion for generative image super-resolution. The code will be released upon publication.
Comments5 tables, 6 figures