arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39051cs.CV

TSMD:用于稳健视频精彩片段检测的时间-流模态丢弃

TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection

  • Inventec Corporation(英业达公司)
  • National Taiwan University(国立台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

Bo-Yuan Cheng, Kuan-Yu Chen, Po-Han Huang, Jeng-Lin Li, Jian-Jiun Ding

AI总结:

针对多模态视频精彩片段检测中时间与流级缺失的稳健性问题,提出TSMD方法,结合结构化缺失模拟与联合目标,在MoSu和Mr. HiSum数据集上显著提升mAP@15。

AI中文摘要:

现有的多模态视频精彩片段检测器通常假设视觉、音频和文本流持续可用。然而,在实际应用中,输入可能遭受局部帧缺失或完整流中断。我们将这一稳健性挑战沿两个维度进行阐述:时间缺失,即每种模态中帧独立缺失;以及流级缺失,即整个视频中某一模态不可用。此外,我们发现均方误差(MSE)损失与评估指标及精彩片段的峰值驱动特性不一致。因此,我们提出时间-流模态丢弃(TSMD),该方法将结构化缺失模拟与联合目标相结合,该联合目标包括逐点MSE、每视频皮尔逊相关性和面向峰值的RankNet损失项。TSMD有三种变体:时间丢弃、流级丢弃和混合丢弃。在MoSu和Mr. HiSum数据集上,在50%独立时间移除下,TSMD-时间在mAP@15上分别比TripleSumm提高7.06和3.41个百分点,而TSMD-流在完整流移除下表现最佳。TSMD-混合保留了这些互补优势的大部分,并在评估的时间与流级条件下排名最佳或第二。

英文摘要:

Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal missingness, where frames are missing independently in each modality, and stream-level missingness, where one modality is unavailable throughout a video. Moreover, we find that the mean squared error (MSE) loss is misaligned with both the evaluation metrics and the peak-driven nature of highlights. Therefore, we propose Temporal-Stream Modality Dropout (TSMD), which combines structured missingness simulation with a joint objective comprising pointwise MSE, per-video Pearson correlation, and peak-oriented RankNet loss terms. TSMD has three variants: temporal, stream-level, and mixed dropout. On the MoSu and Mr. HiSum datasets, TSMD-Temporal improves mAP@15 by 7.06 and 3.41 points over TripleSumm under 50% independent temporal removal, whereas TSMD-Stream performs the best under complete-stream removal. TSMD-Mix retains most of these complementary benefits and ranks the best or the second-best across the evaluated temporal and stream-level conditions.

补充信息

↑