基于增强型视觉-语言对齐的临床可信医学图像描述生成研究
Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对医学图像描述生成的临床可信性问题,提出分离训练与推理对齐的框架,结合多编码器、辅助学习及MedPAIR-SCST方法,通过选择与强化学习对齐提升临床一致性,在数据受限场景下改善医学图像描述的可信度。
AI中文摘要:
医学图像描述生成是一项可加速早期诊断工作流、提升医学诊断AI系统可解释性的技术。然而与通用图像描述生成不同,基于灰度模态、细微解剖线索、专业医学表述及数据质量差异,临床可靠的描述生成仍具挑战性。尽管大型视觉-语言模型近年取得进展,流畅输出未必能保证与临床概念空间或评估标准的充分对齐。为解决该问题,本文提出一种框架,通过分离并增强训练时对齐与推理时对齐来强化临床对齐。构建的医学图像描述生成流水线整合了基于BioMedCLIP和SigLIP2的单/双视觉编码器、Q-Former及基于LLaMA的解码器,并探究了UMLS概念/类型预测的辅助学习贡献。推理时,应用基于单嵌入的重排序从候选中选择最佳描述;训练时,引入MedPAIR-SCST,结合临床相关奖励以将生成分布转向更优临床对齐。实验表明,多编码器设计的互补视觉表征与概念级辅助学习有助于保留临床有意义信息;推理时重排序无需额外训练即可改善语义与临床对齐,而MedPAIR-SCST超越选择,直接优化模型分布以生成更一致、更贴合临床的描述。这些发现显示,联合利用基于选择的对齐与基于强化学习的对齐,即使在数据受限场景下也能推动更可信的医学图像描述生成。
英文摘要:
Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems. However, unlike general image captioning, clinically reliable captioning remains challenging due to grayscale-based modalities, subtle anatomical cues, specialized medical phrasing, and variations in data quality. Despite recent advances in large vision-language models, fluent outputs do not necessarily guarantee sufficient alignment with clinical concept spaces or evaluation criteria. To address this issue, we propose a framework that strengthens clinical alignment by separating and enhancing training-time alignment and inference-time alignment. We build a medical image captioning pipeline that integrates single/dual vision encoders based on BioMedCLIP and SigLIP2, a Q-Former, and a LLaMA-based decoder, and examine the contribution of auxiliary learning for UMLS concept/type prediction. At inference, we apply single-embedding-based reranking to select the best caption among candidates, while at training we introduce MedPAIR-SCST, which combines clinically relevant rewards to shift the generative distribution toward improved clinical alignment. Our experiments show that complementary visual representations with a multi-encoder design and concept-level auxiliary learning help preserve clinically meaningful information. Furthermore, inference-time reranking provides a practical way to improve semantic and clinical alignment without additional training, whereas MedPAIR-SCST goes beyond selection by directly improving the model's distribution to generate more consistent and clinically grounded captions. These findings suggest that jointly leveraging selection-based alignment and reinforcement-learning-based alignment can promote more trustworthy medical image captioning even in data-constrained settings.