arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MR-IQA-2:通过细粒度信用分配实现真实的图像质量反映

MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment

Yuan li, Youyuan Lin, Chenhui Chu, Shin'ya Nishida

arXiv 2608.18579首次发表:更新:

AI 中文总结

本研究针对现有盲图像质量评估(IQA)中推理真实性不足的问题,提出MR-IQA-2框架,通过解耦推理与评分的细粒度信用分配,在IQA基准上实现了与人类评分的竞争性对齐,还能提供更丰富的视觉理解。

AI 中文摘要

多模态大语言模型(MLLMs)在图像质量评估(IQA)领域展现出强大潜力,可提升质量评分与其底层推理的一致性。然而,多数方法通过人类提供的评分监督推理,极少探究其是否真实反映图像质量。仅评分准确性无法保证推理的真实性;共享奖励会模糊监督来源,且当评分偶然正确时,可能强化不真实的推理。为提升盲IQA的真实性与可靠性,本研究旨在:(1)解耦推理与评分的信用分配;(2)为真实推理提供可验证的监督。我们提出MR-IQA-2,这是一个实施推理-编辑-反映流程的actor-editor-judge框架:actor为输入图像生成质量推理,editor依据识别出的质量因子修改图像,冻结的judge对比原始与编辑后的图像,为actor的推理提供反映式监督。MR-IQA-2进一步采用细粒度信用分配解耦推理与评分监督:judge反馈监督推理,人类评分监督预测评分;掩码标记特定更新区分这些信号,同时保留从推理到评分的因果关系。在多个IQA基准测试中,MR-IQA-2实现了与人类的竞争性评分对齐;视觉反映还能提供超越评分的更丰富、更真实的视觉理解,可为图像质量优化及相关下游任务提供参考。代码可在指定URL获取。

英文摘要

Multimodal large language models (MLLMs) have shown strong potential for image quality assessment (IQA) by improving consistency between quality ratings and their underlying reasoning. However, most approaches supervise reasoning through human-provided ratings and rarely examine whether it faithfully reflects image quality. Rating accuracy alone does not ensure faithful reasoning; a shared reward also obscures supervision sources and may reinforce unfaithful reasoning when a correct rating occurs by chance. To improve the faithfulness and reliability of blind IQA, we aim to (1) decouple credit assignment for reasoning and rating and (2) provide verifiable supervision for faithful reasoning. We introduce MR-IQA-2, an actor-editor-judge framework that operationalizes reasoning-editing-reflection. The actor generates quality reasoning for an input image, and the editor revises the image according to the identified quality factors. A frozen judge compares the original and edited images and provides reflective supervision for the actor's reasoning. MR-IQA-2 further uses fine-grained credit assignment to decouple reasoning and rating supervision. Judge feedback supervises reasoning, whereas human ratings supervise the predicted rating. Masked token-specific updates distinguish these signals while preserving the causal relation from reasoning to rating. Across IQA benchmarks, MR-IQA-2 achieves competitive rating alignment with humans. Visual reflection also enables richer and more faithful visual understanding beyond rating, which may inform image-quality optimization and related downstream tasks. Code is available at https://github.com/RobinY99/MR-IQA-2.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑