arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07550cs.CVcs.LG

胸部X光医学视觉语言模型审计:估算跨机构参考一致性

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

  • University College Dublin(都柏林大学学院)
  • School of Computer Science, University College Dublin(都柏林大学学院计算机科学学院)
  • The Third Affiliated Hospital of Southern Medical University(南方医科大学第三附属医院)
  • Dublin City University(都柏林城市大学)

机构由 AI 辅助整理,请以论文原文为准。

Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi

AI总结:

该研究通过评估三个生成式视觉语言模型在胸部X光数据上的表现,估算跨机构参考一致性,发现无法推荐默认估算器,且参考一致性需按站点和接口重新评估。

AI中文摘要:

视觉语言模型通过不暴露置信度分数的接口返回结构化胸部X光检查结果,因此接收机构无法直接判断单个判断的可信度。与某机构参考标准的一致性是否能跨站点、结果类型、预测方向和问题格式转移,在很大程度上尚未被测量。我们在三个机构胸部X光语料库、六个结果类型和两种提示协议下评估了三个生成式视觉语言模型,共包含超过345000个结果级预测,并从少量本地标签预算中估算接收机构的按结果-方向划分的参考一致性。随后在严格的机构留存评估下对估算策略进行压力测试。在完全将接收机构排除在开发之外的评估中,我们设计的七种估算器中的自适应选择并未优于简单的固定替代方案:其平均Brier得分为0.1083,而始终使用Beta-Binomial经验贝叶斯估算器的得分为0.0853,仅针对目标的逻辑回归模型得分为0.0855。这两者相差0.0003,小于该系列对求解器版本变更的敏感性,且在约一半的场景中各有优势,因此无法推荐默认方案。它们相对于跨机构合并估算器的优势集中在一个站点,按机构聚类后无法确认;名义95%水平的插件经验贝叶斯后验预测计数区间覆盖率为87.0%,在最难的机构中更低。因此,参考一致性必须按每个站点和每个接口重新评估;这些结果涉及与机构标签的一致性,而非临床正确性。

英文摘要:

Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an institution's reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured. We evaluated three generative vision-language models on three institutional chest-radiograph corpora and six findings under two elicitation protocols, comprising more than 345,000 finding-level predictions, and estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels. Estimation strategies were then stress-tested under repeated strict institution-held-out evaluation. Under evaluation excluding the receiving institution from development entirely, adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta-Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Those two differ by 0.0003, less than this family's own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness.

补充信息

↑