arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TBSG-Net:用于细粒度视频时刻检索的时间二分场景图网络

TBSG-Net: Temporal Bipartite Scene Graph Network for Fine-Grained Video Moment Retrieval

Ji Huang, Yongsheng Dai, Tianyu Ren, Barry Devereux, Hui Wang

arXiv 2608.02056首次发表:更新:

发表机构

Queen’s University Belfast; School of Electronics, Electrical Engineering and Computer Science(贝尔法斯特女王大学; 电子、电气工程与计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无提案视频时刻检索中静态场景图的时间动态性缺失与时间跨度编码不足问题,提出首个基于动态场景图的TBSG-Net,通过TBSG构造器与混合编码器实现显著性能提升。

AI 中文摘要

近期,无提案视频时刻检索(Video Moment Retrieval, VMR)领域的进展凸显了静态场景图(Static Scene Graphs, SSGs)的有效性。SSGs通过在帧级别建模对象及其关系,丰富了面向检索的视频表示。然而,将SSGs集成到VMR中仍受限于两个固有缺陷:(1)缺乏时间动态性,SSGs无法建模对象及其关系随时间的演变,导致视频表示中关键时间依赖关系的丢失;(2)缺乏显式时间跨度编码,SSGs未对关系的持续时间进行显式编码,使得精确定位颇具挑战性。为解决这些缺陷,我们提出了时间二分场景图网络(Temporal Bipartite Scene Graph Network, TBSG-Net)——据我们所知,这是首个基于动态场景图(Dynamic Scene Graph, DSG)的无提案VMR模型。具体而言,TBSG-Net利用DSGs提取输入视频的以事件为中心的图表示,能够建模对象随时间的交互,从而解决缺陷(1)。这些DSGs随后由新型动态场景图嵌入(Dynamic Scene Graph Embedding, DSG-E)模块处理,以捕获时间跨度和时空信息。首先,DSG-E利用TBSG构造器将DSGs转换为TBSGs,显式编码对象、关系和时间跨度,以解决缺陷(2)。其次,生成的TBSGs被输入混合TBSG编码器,该编码器整合了用于全局事件建模的Transformer变体和用于详细关系推理的图卷积网络,最终生成更全面的时空表示。我们的实验表明,TBSG-Net相比所有基准方法均实现了显著提升。

英文摘要

Recent advances in proposal-free Video Moment Retrieval (VMR) have highlighted the effectiveness of Static Scene Graphs (SSGs). By modeling objects and their relations at the frame level, SSGs enrich retrieval-oriented video representations. However, integrating SSGs into VMR remains constrained by two inherent limitations: (1) Lack of Temporal Dynamics. SSGs fail to model how objects and their relationships evolve over time, leading to the loss of essential temporal dependencies in video representation; and (2) Lack of Explicit Temporal Span Encoding. SSGs do not explicitly encode the duration of relationships, making precise localization challenging. To address these limitations, we propose Temporal Bipartite Scene Graph Network (TBSG-Net)---to the best of our knowledge, the first Dynamic Scene Graph (DSG) based proposal-free VMR model. Specifically, TBSG-Net leverages DSGs to extract event-centric graph representations of the input video, enabling the modeling of object interactions over time and thus addressing limitation (1). These DSGs are then processed by a novel Dynamic Scene Graph Embedding (DSG-E) module to capture both Temporal Span and spatio-temporal information. First, DSG-E utilizes a TBSG Constructor to transform DSGs into TBSGs, explicitly encoding objects, relationships, and time spans to tackle limitation (2). Second, the resultant TBSGs are passed into a hybrid TBSG Encoder that integrates a Transformer variant for global event modeling and a Graph Convolutional Network for detailed relational reasoning, ultimately producing a more comprehensive spatio-temporal representation. Our experiments demonstrate substantial improvements of TBSG-Net over all baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑