SFT后医学图像字幕生成的临床结构化替代奖励
Clinically Structured Surrogate Rewards for Post-SFT Medical Image Captioning
浏览论文内容
中文总结 AI 辅助
该研究针对SFT后医学图像字幕生成,提出临床结构化替代奖励框架,结合生物医学语义等,在ImageCLEFmedical Caption测试集及三个主干上,提升了整体、相关性与事实性指标。
中文摘要 AI 辅助
医学图像字幕生成需要将异质视觉证据转化为简洁的临床描述,尽管表面流畅,但检查结果、断言状态或解剖关系的错误会改变临床意义。序列级策略优化可直接优化完整字幕,但常见奖励依赖全局文本相似度、直接图像-字幕兼容性或无序概念重叠,使视觉邻域和临床声明结构隐含。我们提出一种用于SFT后医学图像字幕生成的临床结构化替代奖励框架,该框架结合生物医学语义和短程词汇保真度,以及两种结构化奖励:分布图像邻域对齐,匹配参考和生成字幕诱导的医学图像库分布;临床图一致性,对实体、断言状态和类型化关系应用最大权重一对一匹配。四种奖励在每个rollout组内独立归一化,结合固定相对权重,通过GDPO优化。在Standard和Synthetical ImageCLEFmedical Caption赛道的组织者评估隐藏测试集及三个视觉-语言主干上,该方法在所有6种主干-赛道组合中,较匹配的SFT基线提升了整体、相关性和事实性,平均相对增益分别为3.4%、2.1%和5.8%。消融实验和配对诊断表明,结构化奖励提供互补信号,减少图像邻域发散并改善实体-断言-关系一致性。
英文摘要
Medical image captioning requires translating heterogeneous visual evidence into concise clinical descriptions, where errors in findings, assertion states, or anatomical relations can alter clinical meaning despite surface-level fluency. Sequence-level policy optimization can directly optimize complete captions, but common rewards rely on global text similarity, direct image-caption compatibility, or unordered concept overlap, leaving visual neighborhoods and clinical-claim structure implicit. We propose a clinically structured surrogate reward framework for post-SFT medical image captioning. The framework combines biomedical semantic and short-range lexical fidelity with two structured rewards: distributional image-neighborhood alignment, which matches the medical-image-bank distributions induced by reference and generated captions, and clinical graph consistency, which applies maximum-weight one-to-one matching to entities, assertion states, and typed relations. The four rewards are independently normalized within each rollout group, combined with fixed relative weights, and optimized with GDPO. Across organizer-evaluated hidden test sets for the Standard and Synthetical ImageCLEFmedical Caption tracks and three vision-language backbones, the method improves Overall, Relevance, and Factuality over matched SFT baselines in all six backbone-track combinations, with average relative gains of 3.4%, 2.1%, and 5.8%, respectively. Ablations and paired diagnostics indicate that the structured rewards provide complementary signals, reducing image-neighborhood divergence and improving entity-assertion-relation consistency.