arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RadPRISM:用于概念解耦图像表示与视觉定位的模式分层放射学报告监督方法

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf, Lena Schmitzer, Lea Schumann, Jannik Kahmann, Friedrich Puttkammer, Johannes Moll, Jannik Lübberstedt, Zeineb Ben Chaaben, Anirudh Narayanan, Cosmin I. Bercea, Sebastian Ziegelmayer, Marcus R. Makowski, Daniel Rueckert, Lisa C. Adams, Keno K. Bressem

arXiv 2608.00147首次发表:更新:

发表机构

Technical University of Munich (TUM); TUM University Hospital; Technical University of Munich, School of Medicine and Health; Klinikum rechts der Isar; Charité – Universitätsmedizin Berlin; Freie Universität Berlin; Humboldt Universität zu Berlin; Imperial College London; Munich Center for Machine Learning (MCML); University Hospital Essen (AöR); Institute for Artificial Intelligence in Medicine (IKIM); Institute of Interventional and Diagnostic Radiology and Neuroradiology; National Center for Tumor Diseases West(慕尼黑工业大学(TUM); 慕尼黑工业大学医院; 慕尼黑工业大学医学与健康学院; 右伊萨尔医院; 柏林夏里特医学院; 柏林自由大学; 柏林洪堡大学; 伦敦帝国理工学院; 慕尼黑机器学习中心(MCML); 埃森大学医院(AöR); 医学人工智能研究所(IKIM); 介入与诊断放射学及神经放射学研究所; 西部肿瘤疾病国家中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RadPRISM将放射学模式作为分层轴,通过专用视觉子空间对齐临床概念,提升零样本分类与视觉定位性能,实现可透明检查的概念解耦医学图像表示。

AI 中文摘要

视觉-语言预训练可从放射学报告中学习丰富的医学图像表示,但以往模型变体通常在单一共享嵌入空间中运行,因此需事后恢复概念级结构与可解释性,这限制了模型透明度及临床实用性。本文提出RadPRISM,将临床医生定义的放射学模式作为指定分层轴:部署的大型语言模型从自由文本报告中提取每个概念的文本片段,每个临床概念在各自专用视觉子空间中对齐,将概念分层转化为直接的顶层对齐监督。该方法在胸部X光片上实例化,采用19概念模式,基于内部多年存档的203602次检查数据,相比匹配的全局对齐基线,RadPRISM将内部数据集零样本分类的宏观AUROC从0.717(95%置信区间0.710-0.723)提升至0.868(95%置信区间0.863-0.872);在外部零样本分类中,其表现与专用的CARZero参考模型相当,但在指游戏视觉定位任务中大幅优于该模型(最高达4.3倍)。此外,放射科医生阅读研究显示,该方法具备概念分层检索能力(前3名的宏观检索正确率为0.78),能呈现报告级检索与固定标签词汇无法表达的解耦描述性发现。RadPRISM可生成具有判别性、空间保真度且原生概念分层的表示,这些表示由临床医生塑造并可被透明检查。

英文摘要

Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility. We introduce RadPRISM, which makes a clinician-defined radiology schema a designated stratification axis: an on-premise large language model extracts per-concept text spans from free-text reports, and each clinical concept is aligned in its own dedicated visual subspace, turning concept stratification into direct, top-level alignment supervision. Instantiated on chest radiographs with a 19-concept schema over $203{,}602$ examinations from an internal multi-year archive, RadPRISM improved internal dataset zero-shot classification from $0.717$ (95% CI, $0.710-0.723$) to $0.868$ (95% CI, $0.863-0.872$) macro AUROC over a matched global-alignment baseline, performed on par with the purpose-built CARZero reference in external zero-shot classification while substantially outperforming it (up to 4.3-fold) in pointing-game visual grounding. In addition, a radiologist reader study demonstrated concept-stratified retrieval ability ($0.78$ macro retrieval correctness rate within rank 3), surfacing disentangled descriptive findings that report-level retrieval and fixed-label vocabularies cannot express. RadPRISM yields discriminative, spatially faithful, natively concept-stratified representations shaped by and transparently inspectable by clinicians.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑