arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29721cs.CVcs.IRcs.MM

SALI:基于电影语法知识的跨镜头关系匹配的镜头感知晚期交互方法

SALI: Shot-Aware Late Interaction for Cross-Shot Relation Matching in Text-to-Video Retrieval using Film-Grammar Knowledge

Toya Oyama, Rainer Lienhart, Shin'ichi Satoh

首次发表
浏览论文内容

中文总结 AI 辅助

针对文本到视频检索中单嵌入丢失人物关系的问题,提出SALI方法,通过提取查询主语宾语并与各镜头嵌入匹配,结合电影语法惩罚,显著提升多镜头关系查询的召回率。

中文摘要 AI 辅助

文本到视频检索通常用单个嵌入表示一个视频片段。这种嵌入往往会丢失人物之间的重要关系。例如,交互“Anna confronts Mark”通常以交替的正反打镜头拍摄(图1a)。单个镜头或对片段所有镜头嵌入取平均无法捕捉这种关系。因此,我们提出SALI(镜头感知晚期交互)。它从单句查询中提取主语和宾语,并将查询、主语和宾语的文本嵌入与视频片段的每个视觉镜头嵌入进行匹配。匹配算子采用贪心最大化或最优传输。微调中的电影语法惩罚项引入微小且一致的偏移。基于CLIP4Clip-meanP构建的SALI在Condensed Movies和ActivityNet上保持整体召回率相当,同时将多镜头关系查询的R@1分别提升3个和12个百分点,在所有对比方法中提升最大,并在MSR-VTT上以整体R@1下降1.4为代价改进了此类查询。

英文摘要

Text-to-video retrieval usually represents a video clip by a single embedding. This embedding often loses important relations between people. E.g., an interaction "Anna confronts Mark" is regularly filmed as alternating shot and reverse shot of both (Fig. 1a). No single shot or averaged embedding over clip shots captures this relation. Thus, we propose SALI (Shot-Aware Late Interaction). It extracts the subject and object from a single-sentence query, and matches the query, its subject and object text embeddings against each visual shot embedding of a video clip. The matching operator is greedy max or optimal transport. A film-grammar penalty in fine-tuning adds a small, consistent shift. Built on CLIP4Clip-meanP, SALI keeps overall recall on par on Condensed Movies and ActivityNet while raising R@1 on multi-shot relation queries by 3 and 12 points, the most among all compared methods, and improves such queries on MSR-VTT at a cost of 1.4 R@1 overall.

发表机构

  • The University of Tokyo(东京大学)
  • National Institute of Informatics(国立情报学研究所)
  • University of Augsburg(奥格斯堡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑