发表机构
University of Sydney; ByteDance; University of Wollongong; University of Adelaide(悉尼大学; 字节跳动; 伍伦贡大学; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对放射学报告生成中报告级监督与疾病级发现的粒度不匹配问题,提出图监督分层临床对齐方法,在3B模型上超越多个更大参数的现有系统,验证了优化监督结构的有效性。
AI 中文摘要
放射学报告生成(Radiology Report Generation, RRG)近期受益于大语言模型,大幅提升了报告流畅度,但临床层面的忠实生成仍具挑战,因为当前监督大多仅在报告层面施加,这造成了粒度不匹配:放射学报告由基于疾病的发现构成,而现有方法主要用整份报告的目标进行训练。为解决该问题,本文提出图监督分层临床对齐(Graph-Supervised Hierarchical Clinical Alignment),将图像-报告监督重新表述为分层临床对齐问题。该方法将此对齐构建为疾病条件化过程,监督分解为两个层级:以疾病为中心的对齐,用于细粒度的疾病特定对应;以及全局临床语义对齐,用于报告层面的语义连贯性。临床知识图谱作为仅训练时使用的结构先验,定义疾病特定监督单元及其临床关系,推理时无额外开销。由于标准对比对齐在研究存在重叠病理时会产生假阴性,本文结合实例条件判别匹配与疾病条件软正则化,实现细粒度且临床一致的跨模态表征。在MIMIC-CXR、IU-Xray和COV-CTR数据集上的实验表明,本文方法在常规指标与临床指标上均持续提升性能,值得注意的是,本文的3B参数模型超越了多个采用更大7B/13B骨干的现有系统,这表明优化监督结构而非增大模型规模,对RRG更为有效。
英文摘要
Radiology report generation (RRG) has recently benefited from large language models, which substantially improve report fluency. However, clinically faithful generation remains challenging because current supervision is still imposed mostly at the report level. This creates a granularity mismatch: radiology reports are composed of disease-grounded findings, while existing methods are trained mainly with whole-report objectives. To address this problem, we propose Graph-Supervised Hierarchical Clinical Alignment, which reformulates image-report supervision as a hierarchical clinical alignment problem. Our method structures this alignment as a disease-conditioned process, where supervision is decomposed into two levels: Disease-Centric Alignment for fine-grained disease-specific correspondence, and Global Clinical Semantic Alignment for report-level semantic coherence. A clinical knowledge graph is used as a training-time-only structural prior that defines disease-specific supervision units and their clinical relationships, introducing no additional overhead at inference. Because standard contrastive alignment could produce false negatives when studies share overlapping pathologies, we combine instance-conditioned discriminative matching with disease-conditioned soft regularization, enabling fine-grained yet clinically consistent cross-modal representations. Experiments on MIMIC-CXR, IU-Xray, and COV-CTR show that our method consistently improves performance on both conventional and clinical metrics. Notably, our 3B model surpasses several prior systems with larger 7B/13B backbones, suggesting that improving supervision structure, rather than increasing model size, can be more effective for RRG.