发表机构
TU Darmstadt; Google Research; TU Munich(达姆施塔特工业大学; 谷歌研究院; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对音频描述(AD)仅关注预分段视频的问题,提出两阶段流程Cue2Narrate,引入LongLSMDC基准,在多段AD生成任务上提升了平均平均精度并优化了AD生成效果。
AI 中文摘要
音频描述(AD)会在对话间隙为视觉事件提供口语叙事,使视障观众能够观看电影。该问题需要确定要叙述的内容(即哪个视觉事件)和叙述时机(即插入AD的位置),以实现最佳用户体验。先前的工作大多将该问题简化为对预分段视频片段的视频字幕,即“内容”在很大程度上是预先确定的,而“时机”则被完全忽略。我们提出了Cue2Narrate,这是一个两阶段流程,可在更长的未修剪电影片段中联合预测要叙述的内容和时机。双模态视听定位器会为每个AD话语预测两个时间上不同的窗口:视觉线索窗口和口语叙事窗口。采用LoRA适配的视觉语言模型(VLM)基于预测的视觉证据生成简洁的AD,其训练使用描述排序损失,该损失将同一帧的字幕(负样本)的排名低于真实AD。为了对这一新问题陈述进行基准测试,我们引入了LongLSMDC基准,包含长达8分钟的电影片段(平均约6.5分钟)。在LongLSMDC上,Cue2Narrate在平均平均精度(avg. mAP)上比仅视觉和仅音频定位基线高出5至12个点。在预测窗口和真实窗口评估下,Cue2Narrate都比对应的微调基础VLM提升了AD生成效果。这些结果确立了首个针对长片段多段AD生成的基准。数据与代码:this https URL
英文摘要
Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visual event) and when (position for inserting the AD) to narrate, to achieve the best user experience. Prior work has largely reduced the problem to video captioning of pre-segmented video clips, i.e., what is largely predefined and when is ignored entirely. We propose Cue2Narrate, a two-stage pipeline that jointly predicts what and when to narrate in longer untrimmed movie clips. A dual-head audio-visual localizer predicts two temporally distinct windows per AD utterance: a visual cue window and a spoken narration window. A LoRA-adapted VLM then generates concise ADs from the predicted visual evidence, trained with a Description Ranking Loss that ranks captions (negative samples) of the same frames lower than the GT AD. To benchmark this new problem statement, we introduce the LongLSMDC benchmark with up to 8-min movie clips (~6.5min on average). On LongLSMDC, Cue2Narrate outperforms video-only and audio-only localization baselines by 5--12 points in avg. mAP. Under both predicted- and GT-window evaluation, Cue2Narrate improves AD generation over the corresponding fine-tuned base VLM. These results establish the first benchmark for multi-segment AD generation on long-form clips. Data & Code: https://github.com/multimodal-ai-lab/Cue2Narrate
CommentsAccepted to EMNLP main conference