arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低接地分数并非未接地的评判者:识别多模态监督中的可感知性混淆

A Low Grounding Score Is Not an Ungrounded Judge: Identifying the Perceptibility Confound in Multimodal Oversight

Rasul Khanbayov, Hasan Kurban

arXiv 2610.00111首次发表:更新:

AI 中文总结

本文提出判决接地分数并揭示其受编辑可感知性混淆影响,导致误报;通过识别缺失量可直接测量误报率,并建议必须搭配未编辑图像的检测探针使用。

AI 中文摘要

模型评判者现在大规模监督多模态系统,过滤训练数据、选择输出,并提供塑造多模态推理模型的奖励。信任一个评判者首先意味着检查它是否使用了其证据,而这一检查本身也值得审视,因此我们提出疑问:视觉接地的反事实探针是否衡量了它所声称衡量的内容。该探针编辑图像使地面真值翻转,保持推理轨迹固定,并询问判决是否随之改变。我们将其形式化为判决接地分数,并表明它不能像此类分数通常被解读的那样被解读。判决仅对触及评判者决策相关阅读的编辑做出响应,因此分数受编辑可感知性的上限约束,除非编辑使属性更易阅读,否则误差是单侧的:分数只能使评判者看起来比实际更不接地。因此,实际失败是误报,即审计者丢弃了一个可用的监督者。在我们陈述的假设下,我们表明这一缺失量不仅是有界的,而且可以从同一审计协议已收集的三个量中识别出来,这使得误报率可直接测量,而不仅仅是一个担忧。审计九个评判者,我们发现预测的排序在我们的整个主要池中严格成立,而那里的典型评判者仅对其属性本可解析的约一半编辑采取行动。应用保守的拒绝阈值可直接证明几个单元为误报,最明显的是一个评判者几乎每次都能检测到注入的错误,同时其分数却仿佛它根本没有使用图像。由此得出的规则是,图像侧反事实分数绝不应单独报告:对未编辑图像的检测探针为其提供上限,证明其误报,且运行成本为零。

英文摘要

Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models. Trusting one means first checking that it uses its evidence, and that check is itself worth scrutinizing, so we ask whether a counterfactual probe of visual grounding measures what it claims to. The probe edits the image so the ground truth flips, holds the reasoning trace fixed, and asks whether the verdict follows. We formalize it as the Verdict Grounding Score and show it cannot be read the way such scores are read. A verdict responds only to an edit that reaches the judge's decision-relevant reading, so the score is capped by how perceptible the edit is, and unless editing makes the attribute easier to read, the error is one-sided: the score can only make a judge look less grounded than it is. The practical failure is therefore a false alarm, an auditor discarding a usable overseer. Under assumptions we state, we show this missing quantity is not merely bounded but identified from three quantities the same audit protocol already collects, which makes the false-alarm rate directly measurable rather than merely a concern. Auditing nine judges, we find the predicted ordering holds strictly across our entire primary pool, and the typical judge there acts on only about half of the edits whose attribute it can otherwise resolve. Applying a conservative rejection threshold certifies several cells as false alarms outright, the clearest being a judge that detects the injected error essentially every time while still scoring as if it had not used the image at all. The rule that follows is that an image-side counterfactual score should never be reported alone: a detection probe on the unedited image upper-bounds it, certifies its false alarms, and costs nothing extra to run.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑