arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13966cs.CVcs.AI

SGWIB:用于视频高光检测的切片Gromov-Wasserstein信息瓶颈

SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection

  • Key Laboratory of Agricultural Machinery Intelligent Control and Manufacturing of Fujian Province University, College of Mechanical and Electrical Engineering, Wuyi University(武夷大学机电工程学院福建省高校农业机械智能控制与制造重点实验室)
  • National Taiwan University of Science and Technology(国立台湾科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Hanjuan Huang, Yung-Chieh Yeh, Hsing-Kuo Pao

中文总结 AI 辅助

提出SGWIB框架,结合结构感知的切片Gromov-Wasserstein信息瓶颈与上下文解耦,提升视频高光检测的片段级预测性能。

中文摘要 AI 辅助

视频高光检测旨在识别视频中捕捉最具信息性或吸引力事件的时间上重要的片段。因此,可靠的预测不仅需要具有判别性的片段表示,还需要保留相邻和远处片段之间的时间关系。信息瓶颈原理已被证明对于学习紧凑且任务相关的表示是有效的,但尚未被探索用于视频高光检测,且直接应用传统公式会忽略片段间的关系结构,并在压缩过程中扭曲高光相关的时间组织。因此,我们引入了切片Gromov-Monge间隙(SGMG),这是一种结构感知的正则化器,用于衡量由指定源到瓶颈映射相对于最优切片结构对应所诱导的过量关系失真。基于SGMG,我们开发了SGWIB,一种用于单模态视频高光检测的信息瓶颈框架,在学习紧凑瓶颈表示的同时保留片段间的时间结构。我们进一步引入了主客场相关上下文伪标签和上下文解耦模块,通过将高光导向信息与上下文模式分离来减少特定于体育的上下文偏差。在MrHiSum和MoSu上的实验表明,SGWIB在两个数据集上均达到了所比较的单模态方法中最佳的Kendall's tau、Spearman's rho、mAP@50和mAP@30。在MrHiSum上,视觉模型在这四个指标上分别将之前的最强结果提高了0.031、0.031、0.87和0.75。这些结果表明,结构感知的信息瓶颈正则化结合上下文解耦改善了片段级别的高光预测。

英文摘要

Video highlight detection aims to identify temporally important segments that capture the most informative or engaging events in a video. Reliable prediction therefore requires not only discriminative segment representations but also preservation of the temporal relationships among neighboring and distant segments. The information bottleneck principle has proven effective for learning compact and task-relevant representations, yet it has not been explored for video highlight detection, and applying conventional formulations directly would overlook inter-segment relational structure and distort highlight relevant temporal organization during compression. We therefore introduce the Sliced Gromov-Monge Gap (SGMG), a structure aware regularizer that measures the excess relational distortion induced by a prescribed source-to-bottleneck mapping relative to an optimal sliced structural correspondence. Building on SGMG, we develop SGWIB, an information-bottleneck framework for single-modal video highlight detection that learns compact bottleneck representations while preserving inter-segment temporal structure. We further introduce Home-Away-Related Contextual Pseudo-Labels and a contextual disentanglement module that reduce sports-specific contextual bias by separating highlight oriented information from contextual patterns. Experiments on MrHiSum and MoSu show that SGWIB attains the best Kendall's tau, Spearman's rho, mAP@50, and mAP@30 among the compared single-modal methods on both datasets. On MrHiSum, the visual model improves the strongest previous results by 0.031, 0.031, 0.87, and 0.75 on these four metrics, respectively. These results show that structure-aware information-bottleneck regularization combined with contextual disentanglement improves segment-level highlight prediction.

补充信息

↑