arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22942cs.CVcs.AI

一种面向开放式图像质量感知的进化式智能体方法

An Evolutionary Agentic Approach for Open-ended Image Quality Perception

Zhenchen Tang, Bo Peng, Zichuan Wang, Songlin Yang, Leilei Cao, Fengjie Zhu, Jing Dong

首次发表
浏览论文内容

中文总结 AI 辅助

提出PACE,一种无需训练的多智能体框架,通过构建显式VQA协议并校准,解决开放式图像质量评估中的整体性偏差,显著降低HOR。

中文摘要 AI 辅助

生成模型正在迅速将图像质量评估(IQA)从传统的保真度因素扩展到新兴维度,如物理合理性和文本渲染正确性。然而,现有的IQA模型依赖于固定定义和重度监督,难以扩展到开放式的感知维度。我们识别出整体性偏差是一个重要限制:在对未见过的维度进行评分时,模型会复用通用的质量先验,导致评分错误和排名反转。为解决这一问题,我们提出了PACE(感知智能体协作进化),一种无需训练的多智能体框架,将开放式IQA表述为显式协议构建。给定目标维度,PACE使用协作智能体构建由可验证的视觉问答(VQA)探针组成的评估协议,将评估基于具体的视觉证据而非整体印象。所得协议仅使用每个维度四张人工标注图像进行校准,同时双轨评分机制使模型感知与人类评分尺度对齐。在传统IQA、结构保真度、上下文感知美学以及新定义的开放式维度上,PACE持续改进其MLLM骨干,在多样化的IQA设置中取得竞争性性能,并将整体覆盖率(HOR)从44.4%降至8.6%。

英文摘要

Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perceptual dimensions. We identify holistic bias as an important limitation: when scoring an unseen dimension, models reuse generic quality priors, leading to scoring errors and rank inversion. To address this, we propose PACE (Perceptual Agentic Collaborative Evolution), a training-free multi-agent framework that formulates open-ended IQA as explicit protocol construction. Given a target dimension, PACE uses collaborative agents to construct an evaluation protocol composed of verifiable Visual Question Answering (VQA) probes, grounding evaluation in concrete visual evidence rather than holistic impressions. The resulting protocol is calibrated using only four human-annotated images per dimension, while a dual-track scoring mechanism aligns model perception with human scoring scales. Across traditional IQA, structural fidelity, context-aware aesthetics, and newly defined open-ended dimensions, PACE consistently improves its MLLM backbone, achieving competitive performance across diverse IQA settings, and reduces the Holistic Override Rate (HOR) from 44.4\% to 8.6\%.

发表机构

  • New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Transsion(传音)

机构由 AI 辅助整理,请以论文原文为准。

↑