arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RefCaptioner:多参考图像接地的视频字幕生成

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang

arXiv 2607.28509首次发表:更新:

AI 中文总结

本文提出多参考图像接地的视频字幕生成新任务,构建语料库与MRVBench基准,提出RefCaptioner框架,实验证实其性能最优且生成字幕更忠实。

AI 中文摘要

现有视频字幕生成模型能生成视频内容的自然描述,但无法将局部视觉元素明确接地到多张参考图像。本文提出多参考图像接地的视频字幕生成这一新任务,要求生成带短语级参考接地的事实性视频描述,并提出两阶段后训练框架RefCaptioner。该框架结合混合数据监督微调(SFT)与分层覆盖折扣组相对策略优化(GRPO),在保留通用视频字幕生成能力的同时,共同提升参考选择、短语级绑定、干扰项排除及跨参考一致性。为支撑训练,本文构建了包含20000个视频和171354张参考图像的语料库,还推出MRVBench基准,用于评估真实世界及AI生成视频的字幕事实性与多参考接地能力。实验表明,RefCaptioner在开源模型中整体性能最优,在标准视频字幕基准上仍具竞争力;人工评估进一步证实,其生成的字幕更受标注者青睐,能通过开源及专有视频生成器实现更忠实于源的视频重建。

英文摘要

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing $20,000$ videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.

Commentshttps://github.com/pkucs-Ltf/RefCaptioner

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑