AI 中文总结
本研究审计VLM裁判在文化情境提示下选择图像的行为,发现4B裁判存在严重位置偏差且弱于CLIP基线,而8B裁判表现更优;跨顺序一致性仅对前者有效,强调裁判更换时需重新审计过滤器。
AI 中文摘要
视觉语言模型(VLMs)日益充当裁判角色,从多张生成的图像中挑选最佳者,因此它们的选择决定了用户所见。此类裁判通常通过与人类评分的分数一致性来验证,而非通过其返回的图像。我们将VLM裁判作为决策者进行审计:在300个文化情境提示上,我们将返回的图像与裁判从未见过的人类评分进行比较,并与从相同候选项中随机选择进行比较,且每次决策都在候选项重新排序后重复进行。一个4B参数的裁判勉强胜过随机选择,且不如CLIP相似度基线。它在49%的调用中选择了第一个显示的图像(随机概率为28%),而重新排序在60%的提示上改变了其选择。对于该裁判,跨顺序的一致性具有信息量:在重新排序后仍保持不变的决策远优于随机,而与较弱的第二裁判的一致性则保留了错误的决策。一个8B裁判几乎不存在位置偏差,且优于CLIP,但同样的过滤器对其而言大多丢弃了好的决策。只有当一致性针对裁判的失败模式时才有帮助,因此每当裁判更换时,过滤器必须重新审计。4B裁判在刻板印象评分上的轻微上升,在跨顺序聚合或使用更大裁判时不再可检测到。
英文摘要
Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the returned image with human ratings the judge never sees and with random choice from the same candidates, and repeat every decision with the candidates reordered. A 4B-parameter judge barely beats random and falls short of a CLIP similarity baseline. It picks the first image shown in 49% of calls (chance: 28%), and reordering changes its choice on 60% of prompts. For this judge, agreement across orders is informative: decisions that survive reordering are much better than random, whereas agreement with a weaker second judge keeps the wrong ones. An 8B judge shows almost no position bias and outperforms CLIP, yet for it the same filter mostly discards good decisions. Agreement helps only when it targets the judge's failure mode, so filters must be re-audited whenever the judge changes. The 4B judge's slight rise in stereotype ratings is no longer detectable after aggregating across orders or with the larger judge.
Comments25 pages including appendix. Code and project page: https://github.com/seochan99/JudgeActs ; data: https://huggingface.co/datasets/seochan99/JudgeActs