视觉语言模型(VLMs)在失明或被误导时的表现如何?针对科学图表的VLMs行为评估
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
浏览论文内容
中文总结 AI 辅助
本研究构建科学图表理解基准SciFigBench,提出A-R-I框架评估VLMs在视觉证据缺失或误导时的行为可靠性,发现高感知与推理准确性不代表行为可靠,不同模型表现存在显著差异。
中文摘要 AI 辅助
现有的视觉语言模型(VLM)基准侧重于感知与推理准确性(VLMs描述和推理图像中内容的能力),但对不确定性下的行为可靠性(视觉证据缺失或误导时的表现)关注有限。我们推出SciFigBench,这是一个用于科学图表理解的诊断性VLM基准,可联合评估感知、推理以及不确定性下的行为可靠性。该基准包含250张图表,涵盖三个评估维度的高质量人工标注,总标注时长超600小时。我们还通过图像变换、推理问题、干扰探针、标题偏差探针以及经确认的选择性模糊目标对这些图表进行扩展,生成了超34000个用于压力测试的评估设置。我们进一步提出Admittance-Resistance-Inductance(A-R-I)框架,用于评估模型是否承认证据不足、抵制误导性上下文并从部分信息中谨慎推理。结果显示不同模型间存在显著行为差异:GPT-5.2的描述质量(MQM 91.6)最高,推理准确性(78.4%)较强,但在96%的案例中会生成无法解读的内容;而能力相当的模型Gemini 3.1 Pro(MQM 90.2,推理81.0%)在71%的此类案例中承认不确定性,且获得最强的抵抗得分(0.91)。这些发现表明,仅高感知与推理准确性无法保证行为可靠性,而这一维度对科学工作流中的部署至关重要。
英文摘要
Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures with high-quality human annotations across three evaluation aspects, totaling 600+ hours of annotation effort. We further extend these figures via image transformations, reasoning questions, resistance probes, caption-bias probes, and confirmed selective-blur targets, producing over 34,000 evaluation setups for stress testing. We further propose the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist misleading context, and infer cautiously from partial information. Our results reveal substantial behavioral differences among models. GPT-5.2 achieves the highest description quality (MQM 91.6) with strong reasoning accuracy (78.4%), yet hallucinates unreadable content in 96% of cases, whereas Gemini 3.1 Pro, a comparably capable model (MQM 90.2, reasoning 81.0%), admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). These findings show that high perception and reasoning accuracy alone do not guarantee behavioral reliability, a dimension critical for deployment in scientific workflows.
发表机构
- University of Aberdeen(阿伯丁大学)
- International Institute of Information Technology Hyderabad(海德拉巴国际信息技术学院)
- University of Technology Nuremberg(纽伦堡工业大学)
- University of Southern California(南加利福尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。