arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15635cs.CVcs.AI

ModaLens:在报告条件下的医学视觉语言模型中测量图像敏感性

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria La… 展开作者

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi

首次发表
浏览论文内容

中文总结 AI 辅助

ModaLens 通过配对图像交换审计,量化报告可用性对医学视觉语言模型图像敏感性的影响,发现报告降低敏感性,并验证了该方向在多个模型中的一致性。

中文摘要 AI 辅助

一份放射学报告已经可以回答临床问题,因此很难判断视觉语言模型是否也使用了图像。ModaLens 是一种配对的图像交换审计方法,用于衡量报告的可用性如何改变图像敏感性:在来自 293 名患者的 3,199 个配对 MIMIC-CXR 病例上,对 MedGemma-27B 进行了测试,每个病例包含全部 14 个问题(13 个针对特定发现的问题和 1 个综合问题),每个图像都被替换为来自另一项研究(通常是同一患者的)的图像,而问题和报告保持不变。在明确的回答指令下,模型生成的答案在有报告时于 4.26% 的试验中发生变化,在没有报告时于 20.94% 的试验中发生变化,配对增加 16.7 个百分点(患者聚类 95% 置信区间 15.6 至 17.7),因此在该协议下,报告的可用性降低了图像交换敏感性;原始提示(使用小写首词读出)给出 4.70% 对比 17.07% 的结果,并且替换也改变了二元预测未变化时的连续答案分数。标签源自报告,这限制了对视觉正确性的结论;该方向在另外两个模型系列中得到了重复验证。代码、确切提示和每个数字的运行记录均可在该 https URL 获取。

英文摘要

A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points (patient-clustered 95 percent CI 15.6 to 17.7), so report availability reduces image-swap sensitivity under this protocol; the original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent, and substitutions also move continuous answer scores where the binary prediction does not change. The labels are derived from reports, which limits conclusions about visual correctness; the direction replicates in two further model lineages. Code, the exact prompts and a run record for every number are at https://github.com/criticaldata/MODALENS.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)
  • American International School Vienna(维也纳美国国际学校)
  • Substrate Labs
  • University of North Florida(北佛罗里达大学)
  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • Motork
  • Collingwood School(科林伍德学校)
  • Beth Israel Deaconess Medical Center(贝斯以色列女执事医疗中心)

机构由 AI 辅助整理,请以论文原文为准。

↑