TiTok:用于多片段时间定位的视听大语言模型
TiTok: Audio-Visual LLM for Multi-Segment Temporal Grounding
查看机构详情
- Ewha Womans University(梨花女子大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
TiTok提出视听大语言模型,通过时间令牌交错与多片段奖励优化,解决未修剪视频中多片段定位的计数失准问题,实现最先进性能。
中文摘要 AI 辅助
未修剪视频中的视听多片段定位(AV-MSG)是一个基础问题,它需要基于视听证据进行推理并为查询预测多个片段,但这一问题仍然具有挑战性。仅基于视觉的模型忽视了互补的声学线索,而视听模型往往无法校准事件数量——我们将这一现象称为计数失准。我们提出了TiTok,一种视听大语言模型(AV-LLM),能够为每个查询定位任意数量的事件时间片段。为了实现精确的边界预测,我们引入了时间令牌交错(TTI)方法,该方法将特殊的时间令牌显式注入视听流中,以对齐输入侧的时间感知与输出侧的时间预测。我们进一步提出了用于强化学习的解耦、面向多片段的奖励,包括全局、局部、计数、精度和格式奖励,并使用组奖励解耦归一化策略优化(GDPO)进行优化。为了评估AV-MSG上的性能,我们建立了一个基于UnAV-100的新评估协议,并提出了CountF1指标,用于量化重叠指标无法捕捉的计数失准。TiTok达到了65.7的mIoU和0.58的CountF1,实现了最先进的性能。我们的代码可在以下链接获取。
英文摘要
Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging. Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration. We present TiTok, an audio-visual large language model (AV-LLM) that localizes an arbitrary number of temporal event segments for each query. For precise boundary prediction, we introduce the Time Token Interleaving (TTI) method, which explicitly injects special time tokens into the audio-visual stream to align input-side temporal perception with output-side temporal prediction. We further propose decoupled, multi-segment-oriented rewards for reinforcement learning, consisting of global, local, count, precision, and format rewards, optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO). To assess the performance on AV-MSG, we establish a new UnAV-100-based evaluation protocol, and propose the CountF1 metric for quantifying count miscalibration that overlap metrics fail to capture. TiTok reaches 65.7 mIoU and 0.58 CountF1, achieving state-of-the-art performance. Our code is available at this link.