AI 中文总结
本文提出多参考图像接地的视频字幕生成新任务,构建语料库与MRVBench基准,提出RefCaptioner框架,实验证实其性能最优且生成字幕更忠实。
AI 中文摘要
现有视频字幕生成模型能生成视频内容的自然描述,但无法将局部视觉元素明确接地到多张参考图像。本文提出多参考图像接地的视频字幕生成这一新任务,要求生成带短语级参考接地的事实性视频描述,并提出两阶段后训练框架RefCaptioner。该框架结合混合数据监督微调(SFT)与分层覆盖折扣组相对策略优化(GRPO),在保留通用视频字幕生成能力的同时,共同提升参考选择、短语级绑定、干扰项排除及跨参考一致性。为支撑训练,本文构建了包含20000个视频和171354张参考图像的语料库,还推出MRVBench基准,用于评估真实世界及AI生成视频的字幕事实性与多参考接地能力。实验表明,RefCaptioner在开源模型中整体性能最优,在标准视频字幕基准上仍具竞争力;人工评估进一步证实,其生成的字幕更受标注者青睐,能通过开源及专有视频生成器实现更忠实于源的视频重建。
英文摘要
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing $20,000$ videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.
Commentshttps://github.com/pkucs-Ltf/RefCaptioner