发表机构
Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对弱监督密集视频字幕中LLM合成过渡字幕缺乏视觉基础的问题,提出SBS框架,利用VLM检测过渡并优化时间掩码,在ActivityNet Captions和YouCook2上取得最优性能。
AI 中文摘要
弱监督密集视频字幕旨在仅给定每个视频的一组有序事件级字幕的情况下,对未修剪视频中的多个事件进行定位和描述。近期研究通过LLM合成辅助过渡字幕,以提供额外的视觉-语言对齐,但这些字幕缺乏视觉基础,且被固定位置和时长生硬地分配给每个事件间间隙。为解决该问题,我们提出Seeing Before Synthesizing(SBS)框架,仅在需要的地方自适应提供有视觉基础的语言引导。利用VLM,我们为事件间间隙生成帧级叙事,并通过这些间隙的语义变化检测过渡;对于已识别的过渡,我们通过融合时间中点与语义变化点,并选择使视觉-语言对齐最大化的宽度来优化事件间时间掩码。在ActivityNet Captions和YouCook2上的实验表明,该方法在字幕和定位任务上均达到了当前最优性能。
英文摘要
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal midpoint with the semantic change point and selecting the width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization.
CommentsAccepted to EMNLP 2026 (main, long)