Re³Cap:基于强化学习的图像字幕增强的检索引导优化
Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
- Taobao & Tmall Group of Alibaba(阿里巴巴淘宝和天猫集团)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Re³Cap方法,利用多模态检索作为推理信号,通过CRS和CQA优化图像字幕,在COCO-LN500基准上关系推理性能较GRPO提升8.64%,优于SFT。
AI中文摘要:
强化学习(RL)在图像字幕任务中已取得显著进展,但在鼓励大型视觉语言模型(LVLMs)探索新颖推理策略方面仍存在局限,这导致RL与监督微调(SFT)之间存在性能差距。本文提出多模态检索可作为字幕优化的有效推理信号,基于此提出Re³Cap(Retrieval-Guided Refinement for Image Captioning),这是一种无需额外标注即可增强图像字幕的检索引导推理策略,由字幕优化建议器(CRS)和字幕质量评估器(CQA)实例化,该策略可识别图像字幕中的幻觉和遗漏,生成更准确、详细的描述。大量实验表明,本文方法在图像字幕任务中优于SFT,尤其在COCO-LN500基准测试中,Re³Cap在关系推理任务上的平均性能比GRPO提升了8.64%。
英文摘要:
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads to a performance gap between RL and Supervised Fine-Tuning (SFT). In this paper, we argue that multi-modal retrieval can serve as an effective reasoning signal for caption refinement. Based on this insight, we present the Retrieval-Guided Refinement for Image Captioning (Re$^3$Cap), a retrieval-guided reasoning strategy that enhances image captioning without requiring additional annotations. Instantiated by Caption Refinement Suggester (CRS) and Caption Quality Assessor (CQA), this strategy identifies hallucinations and omissions in image captions, leading to more accurate and detailed descriptions. Extensive experiments demonstrate the superiority of our method in image captioning, even compared with Supervised Fine-Tuning. Especially, Re$^3$Cap outperforms GRPO with an average improvement of 8.64% in relation reasoning on the COCO-LN500 benchmark.