发表机构
Texas State University(德克萨斯州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对放射学报告生成中的发现级不准确问题,提出位置感知门控和解码器级监督对比损失,在MIMIC-CXR和IU X-Ray上提升了临床效能F1。
AI 中文摘要
放射学报告生成模型能够生成流畅的文本,但仍存在发现级别的错误。诊断驱动的方法通过基于预测的发现来改进生成,但这些预测不会直接修改提供给解码器的视觉补丁特征,且全局门控在空间位置上应用相同的调制。我们引入了一种位置感知门控(PAG),它利用预测的发现表示在空间上调制视觉补丁,无需区域监督。我们还提出了一种解码器级监督对比损失(DSCL),利用共享的阳性发现而非实例身份来构建解码器表示。在MIMIC-CXR上,PAG+DSCL将临床效能(CE)F1从0.484提高到0.502,优于匹配的全局门控参考,而PAG和DSCL单独分别达到0.491和0.495。在不进行额外微调的情况下,组合模型在IU X-Ray上达到0.226的CE F1,而PromptMRG报告的为0.211。
英文摘要
Radiology report generation models can produce fluent text while still containing finding-level inaccuracies. Diagnosis-driven methods improve generation by conditioning on predicted findings, but these predictions do not directly modify the visual patch features provided to the decoder, and global gating applies the same modulation across spatial locations. We introduce a Position-Aware Gate (PAG) that uses predicted finding representations to modulate visual patches spatially without region supervision. We also propose a Decoder-level Supervised Contrastive Loss (DSCL) that structures decoder representations using shared positive findings rather than instance identity. On MIMIC-CXR, PAG+DSCL improves clinical efficacy (CE) F1 from 0.484 to 0.502 over a matched global-gate reference, while PAG and DSCL individually reach 0.491 and 0.495. Without additional fine-tuning, the combined model achieves 0.226 CE F1 on IU X-Ray, compared with 0.211 reported by PromptMRG.
Comments5 pages. Manuscript submitted to ICASSP 2027