发表机构
Lomonosov Moscow State University; MSU Institute for AI, Lomonosov Moscow State University(莫斯科国立大学; 莫斯科国立大学人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散超分辨率图像质量评估,提出DISRQAD数据集与基准,发现现有指标在扩散输出上一致性显著下降,为开发扩散感知质量模型奠定基础。
AI 中文摘要
基于扩散的图像超分辨率(SR)能够生成低分辨率输入所不支持的视觉上合理的细节。我们引入了DISRQAD,一个针对该场景的主观质量数据集和诊断基准。它包含来自十种扩散方法和四种非扩散方法的14,000个SR输出的平均意见分数(MOS),涵盖四种低分辨率退化条件和x2/x4放大倍数。我们评估了51种标准的全参考和无参考指标配置以及11种改编变体。在扩散输出上,与MOS的一致性显著较弱:最强的标准无参考基线在扩散SR上达到0.431 SRCC,而在非扩散SR上为0.813。作为基准使用的案例研究,一个剪枝和蒸馏的Q-ReAlign-mini学生模型在扩散SR上达到0.496 SRCC。DISRQAD衡量感知输出质量,而非对输入的忠实度;它能够分析指标在不同生成器系列和输入条件下的行为。我们的研究结果揭示了扩散SR评估中的显著差距,并为开发对扩散特定伪影敏感的质量模型提供了基础。
英文摘要
Diffusion-based image super-resolution (SR) can create visually plausible detail that is not supported by the low-resolution input. We introduce DISRQAD, a subjective-quality dataset and diagnostic benchmark for this setting. It contains mean opinion scores (MOS) for 14,000 SR outputs from ten diffusion and four non-diffusion methods, spanning four low-resolution degradation conditions and x2/x4 upscaling. We evaluate 51 standard full-reference and no-reference metric configurations and 11 adapted variants. Agreement with MOS is substantially weaker on diffusion outputs: the strongest standard no-reference baseline reaches 0.431 SRCC on diffusion SR versus 0.813 on non-diffusion SR. As a case study in benchmark use, a pruned and distilled Q-ReAlign-mini student reaches 0.496 SRCC on diffusion SR. DISRQAD measures perceived output quality, not faithfulness to the input; it enables analysis of metric behavior across generator families and input conditions. Our findings reveal a substantial gap in the assessment of diffusion-based SR and provide a basis for developing quality models sensitive to diffusion-specific artifacts.