arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03261cs.CVcs.CL

MedQA-MM:医学视觉推理背后的捷径

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

Benlu Wang, Yifan Zhang, Jiaqing Yu, Chin Siang Ong, Juncheng Huang, Zhuohao Li, Zhenyu Zhang, Arman Cohan, Hong Yu, Zonghai Yao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对医学多模态MCQ的推理膨胀问题,构建捷径缓解子集MedQA-MM,通过多维度审计与消融实验揭示模型依赖文本等捷径,证明医学图像推理需路径层面证据。

中文摘要 AI 辅助

基准分数仅奖励最终答案,而非解答问题的路径。在医学多模态多项选择题(MCQ)中,这种区分至关重要,因为正确答案可能由预期的图像发现提供支持,也可能由答案措辞中保留的基准线索、非视觉临床文本、可见图像文本、人工注释或设备/上下文人工制品提供支持。我们将由此产生的分数层面的过度解释称为“推理膨胀”。此处的“路径”是指可支持答案选择的可观察输入路径,而非关于模型隐藏认知的断言。在六个医学多模态MCQ数据集上,我们通过提示侧和图像侧审计、模态消融以及保留医学目标和答案键的匹配修复,将候选线索与行为证据分离。在13种配置的开放模型面板中,全输入准确率为62.63%,仅文本和仅选项设置的准确率分别为53.96%和29.71%。去除长度差距、绝对/显眼以及空间/介词线索后,准确率分别下降6.58、3.50和4.77个百分点。我们还构建了MedQA-MM,这是一个包含1000个项目的捷径缓解子集,其中仅文本和仅选项的准确率降至5.21%和12.33%。这并不意味着模型从不使用图像,而是表明医学图像推理主张需要路径层面的证据。

英文摘要

A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

发表机构

  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • Yale University(耶鲁大学)
  • University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校)
  • VA Bedford Healthcare System(贝斯以色列女执事医疗中心贝德福德分部)
  • Qingdao Medical College of Qingdao University(青岛大学青岛医学院)
  • Yale School of Medicine(耶鲁医学院)
  • National University Hospital, Singapore(新加坡国立大学医院)
  • Zhejiang University(浙江大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑