arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GradeTrap:图像中的权威线索在明确指示忽略的情况下仍会转移VLM的判断

GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them

Deep Dessai

arXiv 2609.06058首次发表:更新:

发表机构

The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GradeTrap通过受控评估揭示,即使明确指示忽略,图像中的权威线索(如官方答案密钥)仍显著转移VLM的判断,且效应强于学生答案或通用对照。

AI 中文摘要

随着视觉语言模型(VLM)能力日益增强并被部署在具有重大影响的现实场景中,它们必须独立评估证据,而不是不加批判地服从人类权威。我们引入了GradeTrap,这是一种受控评估,将两种社会线索置于直接冲突中:一个学生答案(应引起谄媚性认同)和一个归属于同伴、教师或官方答案密钥的冲突答案(应引起基于权威的遵从)。模型在明确指示要独立解决并忽略所有学生答案、反馈和评分标记的情况下,产生自由形式的答案。我们在60个合成的现实权衡场景中测试了模型。五个中性试验建立了稳定的模型相对偏好,随后是六种实验线索(包括对照)的三次重复。在Gemini 3.5 Flash-Lite、GPT-5.6 Luna和Claude Haiku 4.5的45项共同交集上,一个通用的第二答案对照产生了5.4%的冲突答案选择。相对于该对照,汇总的项内变化显示没有可靠的同行评审效应,教师评审效应为6.9个百分点,官方密钥效应为19.5个百分点。相比之下,单独显示冲突的学生答案与单独显示学生参考答案相比,仅将选择从2.2%提高到5.2%。因此,官方密钥来源比学生答案或通用第二答案对照更能转移判断,尽管有明确的忽略指令和与官方密钥一起给出的相反学生答案。效应大小在三个模型之间有所不同。

英文摘要

As vision-language models (VLMs) become increasingly capable and are deployed in consequential real-world settings, they must evaluate evidence independently rather than defer uncritically to human authority. We introduce GradeTrap, a controlled evaluation that places two social cues in direct conflict: a student answer, which should attract sycophantic agreement, and a conflicting answer attributed to a peer, teacher, or official answer key, which should attract authority-based deference. Models produce free-form answers while being explicitly instructed to solve independently and ignore all student answers, feedback, and grading marks. We test the models on 60 synthetic real-world trade-off scenarios. Five neutral trials establish a stable model-relative preference, followed by three repetitions of six experimental cues including controls. On the 45-item common intersection across Gemini 3.5 Flash-Lite, GPT-5.6 Luna, and Claude Haiku 4.5, a generic second-answer control yields 5.4% conflicting-answer selection. Relative to that control, pooled within-item changes show no reliable peer-review effect, a 6.9-point teacher-review effect, and a 19.5-point official-key effect. In contrast, a displayed conflicting student answer alone compared to a displayed student reference answer alone only raises selection from 2.2% to 5.2%. Official-key provenance therefore redirects judgements more than a student answer or the generic second-answer control, despite an explicit ignore instruction and an opposing student answer given along with the official key. Effects vary in magnitude across the three models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑