arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多图像胸部X光片分类中的前瞻性临床指征、事后报告泄露和融合设计:患者聚类评估

Prospective clinical indication, post-hoc report leakage, and fusion design in multi-image chest radiograph classification: a patient-clustered evaluation

Kamran Shahid, Muhammad Munwar Iqbal

arXiv 2607.13800首次发表:更新:

AI 中文总结

研究多图像胸部X光片分类,用多种模型比较,含仅图像、仅指征等。评估不同模型在多指标下表现,发现前瞻性指征与报告目标关联大,排列感知融合有竞争力,事后报告文本存在问题,如造成报告标签循环等。

AI 中文摘要

胸部X光片数据集常将多幅图像与临床指征、检查结果及诊断意见相结合,尽管这些输入是在不同护理阶段产生的。我们评估了15000项ReXGradient - 160K研究,包含两幅可读图像和五项CheXbert衍生的报告观察结果。使用冻结的DenseNet - 121和Bio + ClinicalBERT编码器比较仅图像、仅指征、固定顺序多模态、随机交换、DeepSets和SectionGuard - MI模型。仅将检查结果和诊断意见作为事后泄露控制进行评估。模型用五个种子进行训练,通过2000次患者聚类自助重复估计公共测试不确定性。在U - Ones下,主要图像的宏AUROC为0.643,两幅图像为0.694,指征为0.749,普通双图像加指征融合为0.780。SectionGuard - MI的AUROC为0.783,AUPRC为0.260。相对于普通融合,其配对AUROC差异为0.0031,AUPRC差异为0.0289。DeepSets的前瞻性AUROC点估计最高,随机交换融合的前瞻性AUPRC点估计最高且校准更好。完整报告文本的AUROC为0.979,AUPRC为0.836。这些结果表明前瞻性指征与报告衍生目标密切相关,排列感知融合具有竞争力,事后报告文本会造成大量报告标签循环。

英文摘要

Chest radiograph datasets often combine multiple images with Clinical Indication, Findings, and Impression, although these inputs are produced at different stages of care. We evaluated 15,000 ReXGradient-160K studies with two readable images and five CheXbert-derived report observations. Frozen DenseNet-121 and Bio+ClinicalBERT encoders were used to compare image-only, Indication-only, fixed-order multimodal, random-swap, DeepSets, and SectionGuard-MI models. Findings and Impression were evaluated only as post-hoc leakage controls. Models were trained with five seeds, and public-test uncertainty was estimated with 2,000 patient-cluster bootstrap replicates. Under U-Ones, macro AUROC was 0.643 for the primary image, 0.694 for two images, 0.749 for Indication, and 0.780 for ordinary two-image-plus-Indication fusion. SectionGuard-MI achieved AUROC 0.783 and AUPRC 0.260. Relative to ordinary fusion, its paired AUROC difference was 0.0031 (95% CI, -0.0042 to 0.0104; adjusted p=0.374), while its AUPRC difference was 0.0289 (95% CI, 0.0095 to 0.0413; adjusted p=0.004). DeepSets had the highest prospective AUROC point estimate (0.787), and random-swap fusion had the highest prospective AUPRC point estimate (0.265) with better calibration than SectionGuard-MI. Full report text alone reached AUROC 0.979 and AUPRC 0.836; AUROC remained above 0.973 after exact or expanded masking. These results show that prospective Indication is strongly associated with report-derived targets, permutation-aware fusion is competitive, and post-hoc report text creates substantial report-label circularity.

Comments12 pages, 7 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑