arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24118cs.CL

MC-CXR:用于视觉-语言模型中上下文诱导干扰的多上下文胸部X射线基准

MC-CXR: A Multi-Context Chest X-ray Benchmark for Context-Induced Disruption in Vision-Language Models

Junhyeok Lee, Songsoo Kim, Kyu Sung Choi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对视觉-语言模型在临床胸部X射线解读中受上下文干扰的问题,构建了MC-CXR基准,评估多类模型发现文本误导性比视觉误导性更易导致模型转向错误,验证了上下文诱导干扰的存在。

中文摘要 AI 辅助

视觉-语言模型(VLMs)越来越多地应用于临床流程中,在此类场景下,胸部X射线会与检索到的报告、初步笔记或既往影像一同被解读。现有基准仅衡量模型在孤立情况下能否正确作答,却未考量当合理上下文与影像存在冲突时,模型是否仍能保持仅基于影像的正确决策。我们推出多上下文胸部X射线基准(MC-CXR),该基准包含240个病例,扩展为2522个实例,通过配对扰动分离出上下文诱导的干扰。每个病例固定当前影像和目标发现,同时提供匹配的可靠上下文与误导性上下文,涵盖文本和既往胸部X射线,在有可用视觉叠加的情况下也会纳入。MC-CXR定义了三类任务族和两个配对指标:转向错误率和上下文对齐错误率。我们评估了10个VLMs,涵盖开源通用模型、医学领域模型以及闭源系统。仅靠仅基于影像的准确率是不够的。在误导性文本来源中,平均转向率为45.6%-78.1%;在误导性视觉来源中,平均转向率为35.7%-61.7%。在转向的预测中,74.6%与误导性文本标签对齐,而视觉上下文的这一比例为17.6%,存在57.0个百分点的差距(95%置信区间为50.9-62.8)。在标准化直接作答协议下观察到了这种文本-视觉不对称性。该数据集可在PhysioNet上获取。

英文摘要

Vision-language models (VLMs) are increasingly used in clinical pipelines where a chest X-ray is interpreted alongside retrieved reports, preliminary notes, or prior imaging. Existing benchmarks measure whether models answer correctly in isolation, but not whether they preserve a correct image-only decision when plausible context conflicts with the image. We introduce Multi-Context Chest X-ray (MC-CXR), a benchmark of 240 cases expanded into 2,522 instances that isolates context-induced disruption through paired perturbation. Each case fixes the current image and target finding while presenting matched reliable and misleading context across text and prior CXR, with visual overlays where available. MC-CXR defines three task families and two paired metrics, the switch-to-wrong rate and the context-aligned error rate. We evaluate ten VLMs spanning open-source general, medical-domain, and closed-source systems. Image-only accuracy is necessary but insufficient. Mean switch rates range from 45.6-78.1% across misleading textual sources and 35.7-61.7% across misleading visual sources. Among switched predictions, 74.6% align with the misleading label for text versus 17.6% for visual context, a 57.0-point gap (95% CI 50.9-62.8). This text-visual asymmetry is observed under the standardized direct-answer protocol. The dataset is available on PhysioNet.

发表机构

  • Seoul National University College of Medicine(首尔国立大学医学院)
  • Seoul National University Hospital(首尔国立大学医院)
  • Healthcare AI Research Institute, Seoul National University Hospital(首尔国立大学医院医疗人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑