AI 中文总结
针对视觉语言模型弃权能力评估不足的问题,提出VAD-R基准和Rep2Act方法,通过表示到行动的校准提升自发弃权,在多个模型上显著提高准确率。
AI 中文摘要
视觉语言模型(VLM)对无法回答的问题进行弃权(不执行)的能力,与其准确回答可回答问题同等重要。近期,多个基准被提出以评估和改进VLM的弃权能力,但这些基准存在重大局限。首先,样本中常包含图像或问题中的捷径线索,这些线索揭示了可回答性,而明确的“不可回答”选项进一步阻碍了对自发弃权的准确评估。其次,作为训练数据,它们通常仅提供二元标签,缺乏用于更深层监督的细粒度解释。为解决这些局限,我们引入了带有理据的视觉可回答性诊断(VAD-R),这是一个通过捷径过滤和质量验证的两阶段流程构建的基准,以防止可回答性泄漏。每个示例都标注了逐步理据和因果证据缺口标签。在VAD-R上对最先进的开源和闭源VLM进行评估,结果显示自发弃权有限,平均召回率分别仅为11.4%和16.3%。探测分析表明,某些层中的隐藏状态表示能有效区分可回答性,但这种区分未能体现在最终响应中。受此观察启发,我们提出了Rep2Act,一种表示到行动的校准方法,将潜在的可回答性意识转化为明确的弃权决策。Rep2Act将Qwen2.5-VL-3B在VAD-R上的行动准确率从56.67%提升至86.33%,将Qwen2.5-VL-7B从59.33%提升至88.67%。在分布外的TUBench上,Rep2Act仅用3B模型就取得了53.3%的平均F1分数,分别超过闭源的GPT-4 Turbo和GPT-4o达16.2%和1.1%。
英文摘要
The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substantial limitations. First, samples often contain shortcut cues in images or questions that reveal answerability, while an explicit "unanswerable" option further prevents accurate assessment of spontaneous abstention. Second, as training data, they generally provide only binary labels without fine-grained explanations for deeper supervision. To address these limitations, we introduce Visual Answerability Diagnosis with Rationales (VAD-R), a benchmark constructed through a two-stage pipeline of shortcut filtering and quality verification to prevent answerability leakage. Each example is annotated with step-by-step rationales and causal evidence-gap labels. Evaluation of state-of-the-art open- and closed-source VLMs on VAD-R reveals limited spontaneous abstention, with average recall rates of only 11.4% and 16.3%, respectively. Probing analyses show that hidden-state representations in certain layers can effectively distinguish answerability, yet this distinction fails to manifest in final responses. Motivated by this observation, we introduce Rep2Act, a representation-to-action alignment method that translates latent answerability awareness into explicit abstention decisions. Rep2Act improves action accuracy on VAD-R from 56.67% to 86.33% for Qwen2.5-VL-3B and from 59.33% to 88.67% for Qwen2.5-VL-7B. On the out-of-distribution TUBench, Rep2Act achieves an average F1 score of 53.3% with only a 3B model, surpassing the closed-source GPT-4 Turbo and GPT-4o by 16.2% and 1.1%, respectively.