发表机构
Carnegie Mellon University; Allegheny Health Network; Highmark Health; Johns Hopkins University(卡内基梅隆大学; 阿勒格尼健康网络; 海马克健康; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究量化了放射学报告实践变化对AI报告生成评估指标的影响,提出ReRef方法重写参考报告,并发布MIMIC-CXR-Ext-ReRef数据集,证明参考选择可改变模型排名。
AI 中文摘要
放射科医生遵循异质的报告实践。两位放射科医生检查同一图像并识别出相同的临床发现,却可能撰写表面不同的报告,在术语、缩写、格式和详细程度上有所差异。这些报告规范的变化是评估基于AI的放射学报告生成(RRG)模型时一个未被充分认识的障碍,其中机器生成的报告通常根据其与人类生成的参考报告的一致性进行评估。在本文中,我们量化了已建立的评估指标对报告实践变化的敏感性,揭示的影响足以改变模型的排名。我们引入了一个由放射科医生知情的放射学报告实践变化分类法,以及一种方法(ReRef),该方法沿着我们分类法的轴重写参考报告,同时保留临床解释。例如,在比较九个RRG模型在MIMIC-CXR上使用RadCliQ-v1的性能时,压缩参考报告中正常发现的讨论导致Libra从第一名降至第二名,而CheXOne从第三名升至第一名。我们的结果表明,许多当前指标未能将临床解释与对报告实践的遵从性解耦,并且选择准确反映所需报告实践的“正确”参考在实践中可能很重要。为了支持未来的研究,我们发布了MIMIC-CXR-Ext-ReRef,一个由放射科医生验证的120对(原始,替代)参考报告对数据集,源自MIMIC-CXR。
英文摘要
Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology, shorthand, formatting, and level of detail. These variations in reporting norms represent an under-appreciated obstacle in efforts to evaluate AI-based radiology report generation (RRG) models, where machine-generated reports are typically assessed based on their concordance with human-generated references. In this paper, we quantify the sensitivity of established evaluation metrics to variations in reporting practices, revealing impacts large enough to alter the rankings of models. We introduce a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of our taxonomy while preserving clinical interpretation. For instance, when comparing the performance of nine RRG models on MIMIC-CXR using RadCliQ-v1, condensing the discussion of normal findings in the reference reports causes Libra to drop from first to second place while CheXOne rises from third to first. Our results suggest that many current metrics fail to decouple clinical interpretation from conformity to reporting practices and that choosing the ``right'' references that accurately reflect the desired reporting practices can be important in practice. To support future research, we release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 (original, alternative) reference report pairs derived from MIMIC-CXR.
CommentsPreprint