arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AVSCap:为全模态视频字幕编排视听协同

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

Yanghai Wang, Jiahao Wang, Jiafu Tang, Yuanxing Zhang, Zhe Cao, Hanyan Bian, Zijie Zhang, Weiliang Luo, Zhiyu Pan, Zixuan Dong, Jiaheng Liu, Zhaoxiang Zhang

arXiv 2607.12820首次发表:更新:

发表机构

Nanjing University; Kuaishou Technology; Institute of Automation, Chinese Academy of Sciences(南京大学; 快手科技; 中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究全模态视频字幕编排视听协同问题,提出AVSCap框架,构建训练语料库,采用两阶段策略训练字幕生成器,引入新基准,实验表明该模型在非语音音频覆盖率和跨模态绑定方面表现出色,提升了整体性能。

AI 中文摘要

全模态视频字幕不仅仅是将视觉字幕与音频转录相结合:一个有用的字幕必须描述视觉动作、语音、音乐和音效如何共同演变。现有的大型多模态模型在这个关系步骤上常常失败,将音频和视觉流视为松散耦合的观察结果,依赖自动语音识别,并且未充分说明非语音声音及其与视觉事件的联系。我们提出了AVSCap,一个以明确的跨模态事件绑定为中心的视听字幕框架。首先,我们构建了AVSCap-130K,这是一个通过解耦然后融合的管道生成的三模态训练语料库,在组合有基础的全模态字幕之前锚定视觉和声学证据。其次,我们训练了AVSCap-7B,一个具有两阶段策略的7B字幕生成器:监督微调建立基线能力,而样本高效的强化学习使用混合奖励来优化声学完整性和视听协同。我们的缩放分析表明,强化学习比增加监督微调数据带来更大的收益。第三,我们引入了AVSCapBench基准,该基准将字幕分解为视觉、音频和协同事件,并使用细粒度事件召回率对其进行评估。在AVSCapBench和外部基准上的实验表明,AVSCap-7B提高了非语音音频覆盖率和跨模态绑定,在评估的开源模型中提供了最佳的整体性能。

英文摘要

Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speech, music, and sound effects co-evolve. Existing large multimodal models often fail at this relational step, treating audio and visual streams as loosely coupled observations, relying on automatic speech recognition, and under-specifying non-speech sounds and their links to visual events. We present AVSCap, a framework for audio-visual captioning centered on explicit cross-modal event binding. First, we construct AVSCap-130K, a tri-modal training corpus generated by a decoupled-then-fused pipeline that anchors visual and acoustic evidence before composing grounded omni-modal captions. Second, we train AVSCap-7B, a 7B captioner with a two-stage strategy: supervised fine-tuning establishes baseline capabilities, while sample-efficient reinforcement learning uses hybrid rewards to optimize acoustic completeness and audio-visual synergy. Our scaling analysis shows that reinforcement learning brings larger gains than increasing SFT data. Third, we introduce AVSCapBench, a benchmark that decomposes captions into visual, audio, and synergy events and evaluates them with fine-grained event recall. Experiments on AVSCapBench and external benchmarks show that AVSCap-7B improves non-speech audio coverage and cross-modal binding, delivering the best overall performance among evaluated open-source models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑