arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24105cs.CV

DRRG:用于放射报告生成的离散扩散框架

DRRG: A Discrete Diffusion Framework for Radiology Report Generation

Shaoyang Zhoua, Yingshu Li, Yunyi Liu, Lijun Pu, Lingqiao Liu, Lei Wang, Luping Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于离散扩散大语言模型的DRRG框架,将放射报告生成转化为迭代掩码词去噪,在MIMIC-CXR、CheXpert Plus数据集上的多项指标优于对比方法,为放射报告生成提供了有效替代方案。

中文摘要 AI 辅助

目的:自动放射报告生成(RRG)已被广泛研究,以提高报告准确性并减少放射科医生的工作量。大多数现有方法依赖自回归(AR)框架,该框架逐词生成报告且无法修正早期内容,易导致错误传播,且不符合放射报告的迭代细化过程。相比之下,离散扩散大语言模型(DLLM)通过迭代去噪生成文本,天然支持报告细化,但DLLM尚未在RRG领域得到广泛研究。本研究开发并评估了一种用于RRG的离散扩散框架,支持迭代细化而非传统的从左到右自回归解码。材料与方法:我们开发了DRRG,一种基于DLLM的框架,将RRG表述为迭代掩码词去噪。DRRG整合了临床实体感知互补掩码,以提高词监督覆盖范围并强调临床重要实体,同时配备概念条件模块,将图像衍生的临床概念注入视觉表示。我们在MIMIC-CXR和CheXpert Plus上对DRRG进行训练与评估。结果:在MIMIC-CXR上,DRRG的BLEU-4为0.210、CheXpert-F1为0.549、RadGraph-F1为0.281、GREEN为0.360、RaTEScore为0.604,尽管采用了规模小得多的LLM解码器,仍在多数报告指标上优于对比方法;在CheXpert Plus上,DRRG在对比方法中取得最高的BLEU-4(0.119)和CheXpert-F1(0.347)。结论:离散扩散为自回归放射报告生成提供了有效替代方案,支持迭代、双向报告细化;整合临床聚焦掩码与图像衍生概念条件可提升报告质量与临床一致性。

英文摘要

Purpose: Automatic radiology report generation (RRG) has been widely explored to improve reporting accuracy and reduce radiologists' workload. Most existing methods rely on autoregressive (AR) frameworks that generate reports token by token and cannot revise earlier content, making them prone to error propagation and inconsistent with the iterative refinement process of radiological reporting. In contrast, discrete diffusion large language models (DLLMs) generate text through iterative denoising, naturally enabling report refinement. However, DLLMs have not been extensively investigated for RRG. In this study, we developed and evaluated a discrete diffusion framework for RRG that enables iterative refinement rather than conventional left-to-right autoregressive decoding. Materials and methods: We developed DRRG, a DLLM-based framework that formulates RRG as iterative masked-token denoising. DRRG incorporates a clinical-entities-aware complementary mask to improve token supervision coverage and emphasize clinically important entities, together with a concept-conditioning module that injects image-derived clinical concepts into visual representations. DRRG was trained and evaluated on MIMIC-CXR and CheXpert Plus. Results: On MIMIC-CXR, DRRG achieved BLEU-4 of 0.210, CheXpert-F1 of 0.549, RadGraph-F1 of 0.281, GREEN of 0.360, and RaTEScore of 0.604, outperforming the compared methods on most reported metrics, despite employing a substantially smaller LLM decoder. On CheXpert Plus, DRRG achieved the highest BLEU-4 (0.119) and CheXpert-F1 (0.347) among the compared methods. Conclusion: Discrete diffusion provides an effective alternative to autoregressive radiology report generation by enabling iterative, bidirectional report refinement. Incorporating clinically focused masking and image-derived concept conditioning improves report quality and clinical consistency.

发表机构

  • School of Electrical and Computer Engineering, The University of Sydney(悉尼大学电气与计算机工程学院)
  • Department of Radiology, The First Affiliated Hospital of Soochow University(苏州大学附属第一医院放射科)
  • School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院)
  • School of Computing and Information Technology, University of Wollongong(卧龙岗大学计算与信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

↑