发表机构
Northeastern University; Mohamed bin Zayed University of Artificial Intelligence(东北大学; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出EditJudgeBias基准,通过验证质量保持的反事实干预,审计多模态大语言模型裁判在图像编辑评估中的偏差,发现其受无关线索影响且鲁棒性需多维度量。
AI 中文摘要
多模态大语言模型(MLLMs)越来越多地被用作基于指令的图像编辑的自动裁判,并作为模型训练中的奖励信号。然而,系统地审计这些裁判是否受到与编辑质量无关的线索的影响是具有挑战性的,因为视觉干预本身可能改变被评估的质量。因此,只有当干预被验证保持了底层编辑质量时,判断的转变才能归因于偏差。为了应对这一挑战,我们引入了EditJudgeBias,一个具有已验证质量保持的反事实基准,包含1,196个真实编辑样本和跨四个评估位点注入的13种线索。我们使用校准的多模态验证器、对照和人工检查来验证所请求编辑的质量保持。然后,我们沿着三个互补维度审计五个多模态大语言模型裁判:对质量保持线索的不变性、与人类判断的一致性以及成对偏好的稳定性。重要的是,观察到的转变是针对每个裁判自身的零剂量和重新查询噪声基线进行评估的,而不是针对零。实验表明,质量保持线索使每个裁判都超出了其自身的噪声。虚构的多数意见提高了评分,无关的视觉元素比整个图像操作引起更大的转变,交换候选顺序逆转了高达60.9%的成对决策。编辑区域线索也倾向于降低人类一致性。这三个度量对裁判的刻画不同,表明鲁棒性不能通过单一指标来捕捉。
英文摘要
Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may themselves alter the quality being evaluated. A judgment shift can therefore be attributed to bias only when the intervention is verified to preserve the underlying editing quality. To address this challenge, we introduce EditJudgeBias, a counterfactual benchmark with verified quality preservation, comprising 1,196 real editing samples and 13 cues injected across four evaluation sites. We verify quality preservation for the requested edit using calibrated multimodal validators, controls, and human inspection. We then audit five MLLM judges along three complementary dimensions: invariance to quality-preserving cues, agreement with human judgments, and stability of pairwise preferences. Importantly, observed shifts are evaluated against each judge's own zero-dose and re-query noise floors rather than against zero. Experiments show that quality-preserving cues move every judge beyond its own noise. Fabricated majority opinions increase ratings, irrelevant visual elements cause larger shifts than whole-image manipulations, and swapping candidate order reverses up to 60.9% of pairwise decisions. Edit-region cues also tend to reduce human agreement. The three measures characterize judges differently, showing that robustness cannot be captured by a single metric.
Comments30 pages, 9 figures