发表机构
Macao Polytechnic University; Netherlands Cancer Institute; Radboud University Medical Center(澳门理工大学; 荷兰癌症研究所; 拉德堡德大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估系统一模型Jev作为放射学报告事实性评判器的性能,通过双向支持判断检测遗漏与无依据主张,在多个基准上优于开放NLI评判器,且成本低、准确性高。
AI 中文摘要
AI生成的放射学报告可能看起来与医生撰写的报告相似,却遗漏了异常、添加了无依据的发现或颠倒了其存在性。衡量这些事实性差异对于评估报告生成器至关重要。我们研究了Jev,一个系统一决策模型,作为与医生撰写的参考报告一致性的简单、低成本评判器。我们的评估器检查每个陈述是否得到另一份报告的支持,并双向结合这些判断以捕捉无依据的主张和遗漏。在RadEvalX上,单问题配置达到Kendall相关系数0.573,在RadEvalExpert上与专家错误计数的相关系数为0.398,在匹配的分解和聚合下优于开放的自然语言推理评判器。每个陈述一个支持问题在保持与七个问题相似的专家一致性的同时,使用的判断输入令牌减少了43-45%。按文档记录的API价格,每百对报告对的判断成本低于三美分,不包括本地分解。在单独的受控错误测试中,Jev检测错误否定的AUROC为0.977。本地RadMatch在专家数据集中的临床显著错误以及共享RadEvalExpert子集中的总错误上取得了更强的一致性。发现计数和错误范围分析表明,基准一致性反映了报告大小和错误定义以及医学错误检测。这些结果支持Jev作为衡量生成放射学报告中事实性差异的实用判断组件,并指出了更精细评估仍然有价值的领域。
英文摘要
An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.