arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10505cs.AIcs.CLcs.CV

RadFusion:迈向阈值可控的放射学报告生成

RadFusion: Towards Threshold-Controllable Radiology Report Generation

Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz

首次发表
浏览论文内容

中文总结 AI 辅助

RadFusion框架为放射学报告生成赋予阈值可控性,融合多标签分类器与VQA报告生成器,提升诊断准确率,支持ROC分析,适配不同临床场景,更易通过监管审批。

中文摘要 AI 辅助

自动化放射学报告生成正因放射科医生短缺而快速发展,但与感知模型不同,现有生成模型无法对其诊断内容的灵敏度-特异度权衡进行控制。这种控制至关重要,因为临床场景存在差异:急诊分诊优先考虑灵敏度以减少漏诊,而确诊解读则强调特异度以限制不必要的干预。单一固定的报告既无法适应这些场景,也无法支持监管审批广泛要求的基于ROC的验证。我们提出RadFusion,这一框架为报告生成赋予阈值可控性。我们的方法将提供每种疾病置信分数的多标签分类器,与基于VQA的、详细描述医学发现的报告生成器相融合;随后大语言模型(LLM)重写报告,使其中所述诊断符合所选阈值下分类器的决策,同时基于生成器的描述。在MIMIC-CXR数据集上,RadFusion的性能符合分类器的ROC曲线:扫描阈值并将报告映射回类别标签可复现分类器已验证的ROC性能。这种一致性使得生成的报告可通过ROC分析进行定量评估,为监管审批提供了有力支撑,并支持选择与临床场景匹配的操作点。此外,结合两种模型类型的方法比无控制生成的诊断准确率更高:在匹配特异度时灵敏度提升6.9%,在匹配灵敏度时特异度提升20.7%。这些结果表明,RadFusion使报告生成具备临床适应性、可定量验证性,且诊断可靠性更高。

英文摘要

Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions. A single fixed report can neither adapt to these scenarios nor support the ROC-based validation widely expected for regulatory clearance. We introduce RadFusion, a framework that equips report generation with threshold controllability. Our method fuses a multi-label classifier, which provides per-disease confidence scores, with a VQA-based report generator, which describes medical findings in detail; an LLM then rewrites the report so that its stated diagnoses follow the classifier's decisions at the selected threshold while staying grounded in the generator's descriptions. On MIMIC-CXR, the performance of RadFusion conforms to the classifier's ROC curve: sweeping the threshold and mapping the reports back to class labels reproduces the classifier's validated ROC performance. This conformance makes generated reports quantitatively evaluable through ROC analysis, strengthening the case for regulatory clearance, and enables operating-point selection that matches report behavior to clinical context. Moreover, combining the two model types improves diagnostic accuracy over uncontrolled generation: sensitivity increases by 6.9% at matched specificity, and specificity by 20.7% at matched sensitivity. These results show that RadFusion makes report generation clinically adaptable, quantitatively verifiable, and diagnostically more reliable.

发表机构

  • Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

↑