arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.09393cs.CV

CapRL++:基于可验证奖励的统一强化学习用于密集图像和视频描述生成

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

Penghui Yang, Long Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Yibin Wang, Yujie Zhou, Jiazi Bu, Jianze Liang, Qidong Huang, Jiaqi Wang, Feng Wu, Dahua Lin

首次发表 更新
浏览论文内容

中文总结 AI 辅助

提出CapRL++框架,利用可验证奖励的强化学习(RLVR)优化多模态描述生成,通过非视觉语言模型回答问题的准确性作为奖励,提升密集描述质量,在20多个基准上超越传统监督微调。

中文摘要 AI 辅助

图像和视频描述是连接视觉与语言领域的基础任务,在预训练大型视觉语言模型(LVLMs)中发挥关键作用。当前最先进的描述模型通常采用监督微调(SFT)训练,这种范式依赖于昂贵且不可扩展的标注,并常导致模型记忆特定真实答案,限制了其通用性和生成多样化、创造性描述的能力。为克服这些局限,我们提出将可验证奖励的强化学习(RLVR)应用于多模态描述的开放任务。我们引入描述强化学习++(CapRL++),一种新颖的无参考训练框架,通过效用重新定义描述质量:高质量描述应使非视觉语言模型能够准确回答关于相应视觉内容的问题。CapRL++采用解耦的两阶段流程,其中LVLM生成描述,目标奖励来自一个独立的、无视觉的LLM仅基于该描述回答多项选择题的准确率。在超过20个图像和视频基准上的评估表明,CapRL++提升了密集描述质量,并增强了基于描述的预训练在空间和时间理解等任务上的表现。在CapRL++标注的可扩展图像和视频描述数据集上预训练带来了显著的下游收益。此外,在描述质量评估的Prism框架内,使用CapRL++训练的紧凑模型在密集描述性能上可与Qwen2.5-VL-72B和Qwen3-VL-235B-A22B等大得多的模型相媲美。这些结果验证了CapRL++能有效训练模型生成可泛化、高保真的描述,为超越传统SFT的局限奠定了坚实基础。

英文摘要

Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Current state-of-the-art captioning models are typically trained with Supervised Fine-Tuning (SFT), a paradigm that relies on expensive, non-scalable annotations and often causes models to memorize specific ground-truth answers, limiting their generality and ability to generate diverse, creative descriptions. To overcome these limitations, we propose applying Reinforcement Learning with Verifiable Rewards (RLVR) to the open-ended task of multimodal captioning. We introduce Captioning Reinforcement Learning++ (CapRL++), a novel reference-free training framework that redefines caption quality through its utility: a high-quality caption should enable a non-visual language model to accurately answer questions about the corresponding visual content. CapRL++ employs a decoupled two-stage pipeline where an LVLM generates a caption, and the objective reward is derived from the accuracy of a separate, vision-free LLM answering Multiple-Choice Questions based solely on that caption. Evaluations on more than 20 image and video benchmarks show that CapRL++ improves dense caption quality and strengthens caption-based pretraining across tasks such as spatial and temporal understanding. Pretraining on scalable image and video caption datasets annotated by CapRL++ yields substantial downstream gains. Furthermore, within the Prism Framework for caption quality evaluation, compact models trained with CapRL++ achieve dense captioning performance comparable to substantially larger models such as Qwen2.5-VL-72B and Qwen3-VL-235B-A22B. These results validate that CapRL++ effectively trains models to produce generalizable, high-fidelity descriptions, establishing a robust foundation beyond the limitations of traditional SFT.

发表机构

  • Tsinghua University(清华大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Microsoft(微软)
  • Shanghai AI Laboratory(上海人工智能实验室)
  • Shanghai Innovation Institute(上海创新研究院)
  • Alibaba Cloud(阿里云)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑