发表机构
Raidium; Fondation Ophtalmologique Adolphe de Rothschild; Gustave Roussy(Raidium; 阿道夫·德·罗特希尔德眼科基金会; 古斯塔夫·鲁西癌症中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RadMatch是一种基于LLM的多阶段放射学报告评估指标,通过发现级匹配在七个临床属性维度评分,在两个专家基准上临床一致性最优,将开源代码及交互式仪表板。
AI 中文摘要
随着AI系统越来越多地用于起草放射学报告,可靠评估其临床质量仍是一项关键挑战。基于大语言模型(LLM)的指标目前与放射科医生判断的相关性最高,但它们输出单一不透明分数,临床医生或模型构建者均难以解释或审计。我们提出RadMatch,这是一种基于LLM的多阶段指标,它将报告比较分解为结构化的发现级匹配,在七个临床属性维度(状态、位置、严重程度、形态、确定性、纵向对比和测量)上进行感知重要性的评分和错误表征。主要分数为可操作错误计数,兼具可解释性和可审计性。候选发现被评为正确、部分正确或不正确,未匹配的发现则计为遗漏或幻觉。分类、可操作安全召回率/精度以及按子集视图提供了互补的、面向部署的视角。在两个专家基准上,RadMatch是与临床一致性最高的指标,在ReXVal上匹配放射科医生间的一致性,在更具挑战性的RadEvalExpert上的表现超过此前最佳指标一倍以上。仅依靠少样本提示,它被设计为可扩展至其他模态和解剖部位。我们将发布RadMatch的开源代码及用于检查结果的交互式仪表板。
英文摘要
As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge. Large language model (LLM)-based metrics are now the best-correlated with radiologist judgment, yet they output a single opaque score that neither a clinician nor a model builder can easily interpret or audit. We introduce RadMatch, a multi-stage, LLM-based metric that decomposes report comparison into a structured finding-level matching with significance-aware scoring and error characterization across seven clinical attribute dimensions (status, location, severity, morphology, certainty, longitudinal comparison, and measurement). The main score is the actionable-error count, both interpretable and auditable. Candidate findings are graded correct, partial, or incorrect, and unmatched findings are counted as missed or hallucinated. Triage and actionable safety recall/precision and per-subset views add complementary, deployment-oriented lenses. Across two expert benchmarks, RadMatch is the most clinically aligned metric, matching inter-radiologist agreement on ReXVal and more than doubling the best prior metric on the harder RadEvalExpert. Relying only on few-shot prompting, it is designed to extend to other modalities and anatomies. We will release RadMatch as open-source code with an interactive dashboard for inspecting results.
CommentsAccepted to ECCV 2026 Workshop on Medical Foundation Models and Benchmarks