发表机构
University of Waterloo; NVIDIA; National Taiwan University; Nanyang Technological University(滑铁卢大学; 英伟达; 国立台湾大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VIEScore2通过网格表示联合预测图像质量分数和缺陷位置,利用GRPO优化定位,在评估任务上超越通用视觉语言模型,实现可解释的空间化图像评估。
AI 中文摘要
现有的合成图像评估器通常仅提供标量质量分数,且不识别支持该分数的图像区域。我们提出了VIEScore2,一个用于图像生成和编辑任务(可带条件图像)的统一评估器。VIEScore2将图像表示为N×N网格,并在单次模型前向传播中联合预测质量分数和缺陷位置。其基于文本的网格表示提供了异构空间监督的通用接口,并实现了可直接验证的训练后目标。我们在涵盖仅分数、仅定位以及生成和编辑任务中的联合监督的38K个样本上进行训练。从监督微调开始,我们进一步应用GRPO,利用结合单元格级Dice重叠、分数准确性和输出格式有效性的奖励来改进缺陷定位。一个无需参数的解析器将结构化预测转换为可读的解释。在主测试套件上,VIEScore2实现了0.601的总体分数SRCC,而匹配输入下最强的零样本通用视觉语言模型基线Gemini-3-Flash为0.491。对于缺陷定位,在六个基准中的三个上,VIEScore2在每图像网格IoU上优于通用视觉语言模型和专门的空间评估器,并在五个基准上排名前三,其中包括其训练来源之外的基准。
英文摘要
Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.
CommentsPreprint. Project page: https://tiger-ai-lab.github.io/VIEScore2/