学习局部感知并与病理学语义临床对齐的放射学报告生成方法
Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation
浏览论文内容
中文总结 AI 辅助
针对放射学视觉-语言模型的图像-文本对齐缺陷,提出病理学感知对齐框架PALM,结合共享病理学原型与掩码证据建模,在多数据集上提升报告生成与异常鲁棒性
中文摘要 AI 辅助
近期适配放射学的视觉-语言模型在标准报告生成基准上已取得优异性能,但视觉与文本特征间的不完善对齐及相关性仍限制了其鲁棒性与泛化性。现有方法通过自回归报告监督隐式连接图像与文本,或通过对比学习显式连接。然而仅自回归监督不足以建立可靠的图像-文本对齐,而对比学习会将描述相关病理学的未配对报告分开,仅因它们未与同一图像配对。这在放射学中存在问题,因为不同报告可能共享兼容的病理学语义,而非真正的负样本。结果是,学习到的表征可能无法围绕共享病理学概念组织图像与报告,导致解码器依赖预训练语言先验,生成临床看似合理但未完全得到放射学证据支持的报告。为解决此问题,我们提出PALM,一种用于放射学报告生成的病理学感知对齐框架。PALM不直接匹配每个图像-报告对并分开所有其他对,而是通过共享病理学原型对齐视觉与文本特征。这些原型在放射学证据与文本发现间提供临床有意义的桥梁,使具有相似病理学语义的病例能向共同概念靠拢,而不分开兼容病例。此外,我们引入掩码证据建模,通过学习被掩码图像区域导致的语义变化,增强图像编码器对局部放射学证据的敏感性。在MIMIC-CXR、IU X-Ray和MIMIC-ABN上的实验表明,PALM始终提升报告生成及异常聚焦的鲁棒性。
英文摘要
Recent radiology-adapted vision-language models have achieved strong performance on standard report generation benchmarks, yet their robustness and generalization remain constrained by imperfect alignment and correlation between visual and textual features. Existing methods connect image and text either implicitly through autoregressive report supervision or explicitly through contrastive learning. However, autoregressive supervision alone is insufficient to establish reliable image-text alignment, while contrastive learning can push apart unpaired reports that describe related pathologies simply because they are not paired with the same image. This is problematic in radiology, where different reports may share compatible pathology semantics rather than being true negatives. As a result, the learned representation may fail to organize images and reports around shared pathology concepts, causing the decoder to rely on pretrained language priors and generate clinically plausible reports that are not fully supported by radiographic evidence. To address this issue, we propose PALM, a pathology-aware alignment framework for radiology report generation. Instead of directly matching each image-report pair while separating all others, PALM aligns visual and textual features through shared pathology prototypes. These prototypes provide a clinically meaningful bridge between radiographic evidence and textual findings, allowing cases with similar pathology semantics to move toward common concepts without separating compatible cases. In addition, we introduce Masked Evidence Modeling to strengthen the image encoder sensitivity to local radiographic evidence by learning semantic changes caused by masked image regions. Experiments on MIMIC-CXR, IU X-Ray, and MIMIC-ABN show that PALM consistently improves both report generation and abnormality-focused robustness.