发表机构
IIE, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Xiaohongshu Inc.; Nankai University(中国科学院自动化研究所; 中国科学院大学; 小红书科技有限公司; 南开大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GS-IQA框架,将图像质量评估重构为渐进式“何处—什么—如何”诊断,采用两阶段强化学习,构建Diag-Bench基准,在失真定位、识别和严重性估计上超越现有方法。
AI 中文摘要
多模态大语言模型(MLLMs)通过将视觉感知与描述性评估相结合,在图像质量评估(IQA)领域展现出巨大潜力。然而,现有方法主要侧重于整体质量预测,往往作为黑盒运作,对失真发生的位置及其如何影响感知质量提供的洞察有限,阻碍了对局部和异质退化的细粒度分析。我们提出GS-IQA,一个将IQA重构为渐进式“何处—什么—如何”诊断的框架,模拟了从最初一瞥到更仔细审视的人类感知过程。由于严重性判断仅对正确定位和识别的区域才有意义,我们通过一种尊重这种依赖关系的两阶段强化学习范式来实现这一进程:一瞥阶段使用感知门控奖励来确定退化所在位置及其类型,仅在两者都正确时才激活严重性反馈;而细察阶段引入在线奖励条件退化生成,以合成针对模型感知瓶颈的困难样本,增强其对细微严重性差异的辨别能力。为进行系统评估,我们构建了Diag-Bench,一个区域级IQA基准,包含约25K个精选样本,涵盖12种失真类型和五个有序严重性级别。大量实验表明,GS-IQA在失真定位、识别和严重性估计方面持续超越最先进方法,且其诊断表示可有效迁移到跨多个外部基准的常规全局质量预测中。代码和数据将发布。
英文摘要
Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment (IQA) by bridging visual perception with descriptive evaluations. However, existing approaches mainly focus on holistic quality prediction, often functioning as black boxes that provide limited insight into where distortions occur and how they affect perceived quality, hindering fine-grained analysis of localized and heterogeneous degradations. We propose GS-IQA, a framework that reformulates IQA as a progressive Where--What--How diagnosis, emulating the human perceptual process from an initial glance to closer scrutiny. Since a severity judgment is meaningful only for a correctly localized and recognized region, we realize this progression through a two-stage reinforcement learning paradigm that respects such dependencies: the glance stage uses a perception-gated reward to establish where degradations lie and what they are, activating severity feedback only once both are correct, while the scrutiny stage introduces online reward-conditioned degradation generation to synthesize hard examples targeted at the model's perceptual bottlenecks, sharpening its discrimination of subtle severity variations. To enable systematic evaluation, we construct Diag-Bench, a region-level IQA benchmark of about 25K curated samples spanning 12 distortion types and five ordinal severity levels. Extensive experiments show that GS-IQA consistently surpasses state-of-the-art methods in distortion localization, recognition, and severity estimation, and that its diagnostic representations transfer effectively to conventional global quality prediction across diverse external benchmarks. Code and data will be released.