用于准确且可解释的放射学报告生成的诊断链框架
A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- Department of Computer Science and Engineering, Hong Kong University of Science and Technology(计算机科学与工程系,香港科学理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对放射学报告生成临床疗效不佳且缺乏可解释性的问题,提出诊断链框架CoD。该框架通过诊断对话生成QA对提示大语言模型生成准确报告,并设计诊断与病灶定位模块增强可解释性与工作效率,结合全监督学习策略,在多个基准上优于现有模型。
AI中文摘要:
尽管放射学报告生成(RRG)取得了进展,但现有工作仍面临两个挑战:1)临床疗效表现不佳,尤其是在病灶属性描述方面;2)生成的文本缺乏可解释性,使得放射科医生难以信任结果。为了应对这些挑战,我们专注于一个可信赖的RRG模型,它不仅生成准确的异常描述,还提供其预测的依据。为此,我们提出了一个名为诊断链(CoD)的框架,该框架维持一个诊断过程链,以实现临床准确且可解释的RRG。它首先通过诊断对话生成问答(QA)对以提取关键发现,然后利用QA诊断提示大型语言模型进行准确生成。为了增强可解释性,设计了一个诊断定位模块来匹配QA诊断和生成的句子,其中诊断作为参考。此外,设计了一个病灶定位模块来定位图像中的异常,进一步提高放射科医生的工作效率。为了促进标签高效的训练,我们提出了一种具有临床一致性的全监督学习策略,以利用来自不同数据集的各种类型的标注。我们的努力带来了:1)一个带有QA对和病灶框的全标注RRG数据集;2)一个用于评估报告在描述病灶位置和严重性方面准确性的评估工具;3)广泛的实验以证明CoD的有效性,它在两个RRG基准上上一致地优于专家和通才模型,并通过将生成句子准确关联到QA诊断和图像,展现出可期的可解释性。summary_cn
英文摘要:
Despite the progress of radiology report generation (RRG), existing works face two challenges: 1) The performances in clinical efficacy are unsatisfactory, especially for lesion attributes description; 2) the generated text lacks explainability, making it difficult for radiologists to trust the results. To address the challenges, we focus on a trustworthy RRG model, which not only generates accurate descriptions of abnormalities, but also provides basis of its predictions. To this end, we propose a framework named chain of diagnosis (CoD), which maintains a chain of diagnostic process for clinically accurate and explainable RRG. It first generates question-answer (QA) pairs via diagnostic conversation to extract key findings, then prompts a large language model with QA diagnoses for accurate generation. To enhance explainability, a diagnosis grounding module is designed to match QA diagnoses and generated sentences, where the diagnoses act as a reference. Moreover, a lesion grounding module is designed to locate abnormalities in the image, further improving the working efficiency of radiologists. To facilitate label-efficient training, we propose an omni-supervised learning strategy with clinical consistency to leverage various types of annotations from different datasets. Our efforts lead to 1) an omni-labeled RRG dataset with QA pairs and lesion boxes; 2) a evaluation tool for assessing the accuracy of reports in describing lesion location and severity; 3) extensive experiments to demonstrate the effectiveness of CoD, where it outperforms both specialist and generalist models consistently on two RRG benchmarks and shows promising explainability by accurately grounding generated sentences to QA diagnoses and images.