arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05540cs.CVcs.AI

知道何时不回答:视觉-语言模型中的弃权与拒绝推理

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel, Hansa Meghwani, Jyotika Singh, Ranjeet Gupta, Graham Horwood, Tao Sheng, Avi Sil, Sujith Ravi, Dan Roth

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLMs在无法从图像回答的医学诊断查询中的弃权行为,提出PARITY数据集并审计,发现模型间存在拒绝与推测分歧,表情影响选择,临床护栏可提升弃权率。

中文摘要 AI 辅助

许多医学状况需要通过详细、多情境的临床评估来诊断,而非仅凭视觉外观。尽管如此,视觉-语言模型(VLMs)越来越多地被用于以涉及医学或诊断判断的方式解读图像,当此类推断缺乏依据时,这引发了安全隐患。自闭症谱系障碍(ASD)的诊断需要行为和发展证据,而非静态的面部照片。我们审查了VLMs是否会对这一无法回答的配对图像查询进行弃权(不执行),以及表情是否会左右非弃权的选择。我们引入了PARITY(使用复用身份的配对评估),这是一个合成的、人口统计学上平衡的身份控制中性/表情肖像对数据集,并配有中性-中性对照。所有身份均为合成且无ASD状态;由于该查询无法从图像中回答,任何非弃权的选择均被视为有害归因。在当代VLMs中,我们发现了拒绝优先模型与推测性模型之间的明显分野;在后者中,某些表情不成比例地触发了有害选择。临床护栏和单图像框架显著提高了弃权率,提示在提示词和界面设计中存在可操作的缓解措施。

英文摘要

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such inferences are unsupported. ASD diagnosis requires behavioral and developmental evidence, not static facial photographs. We audit whether VLMs abstain from this unanswerable paired-image query, and whether expressions sway non-abstaining choices. We introduce PARITY (Paired Assessment with Reused Identity), a synthetic, demographically balanced set of identity-controlled neutral/expression portrait pairs with neutral-neutral controls. All identities are synthetic and have no ASD status; because the query is unanswerable from images, any non-abstaining selection is treated as a harmful attribution. Across contemporary VLMs, we find a clear split between refusal-first models and speculative models; in the latter, certain expressions disproportionately trigger harmful selections. Clinical guardrails and single-image framing substantially increase abstention, suggesting actionable mitigations in both prompting and interface design

发表机构

  • Oracle AI(甲骨文人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑