用于部分相关视频检索的CLIP固有时间适应
Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval
浏览论文内容
中文总结 AI 辅助
针对PRVR问题,提出ITA框架,通过骨干内部时间适应和亲和度加权梯度传播优化CLIP,在PRVR基准实现SOTA性能,跨数据集迁移稳健且帧级证据更准确。
中文摘要 AI 辅助
部分相关视频检索(PRVR)旨在检索包含与文本查询相关片段的未修剪视频。由于目标片段仅占视频的一部分,PRVR需要基于超出粗粒度视频级匹配的细粒度理解进行检索。然而,现有方法通常依赖冻结的CLIP帧特征,这些特征缺乏时间理解。即使在最近的参数高效CLIP适应进展下,视频级预测仍可能由不准确的帧级证据支撑。本文中,我们提出用于PRVR的固有时间适应(ITA)框架。首先,我们的骨干内部时间适应允许最后几个视觉Transformer层关注相邻帧组,这提供了具有时间感知的帧嵌入,同时保持CLIP冻结且仅训练适应参数。其次,我们引入亲和度加权梯度传播以解决PRVR的弱监督性质,基于文本-帧亲和度柔和聚合前k个帧,并将学习信号传播至多帧查询相关帧。我们的方法在PRVR基准上实现了最先进的性能,展示了稳健的跨数据集迁移,并在真实查询相关片段内检索到显著更准确的帧级证据。我们的代码可在此http URL获取。
英文摘要
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos that contain moments relevant to a text query. Since the target moment occupies only a portion of the video, PRVR requires retrieval based on fine-grained understanding beyond coarse video-level matching. However, existing methods often rely on frozen CLIP frame features, which lack temporal understanding. Even with recent progress in parameter-efficient CLIP adaptation, video-level predictions can still be supported by imprecise frame-level evidence. In this paper, we propose an Intrinsic Temporal Adaptation (ITA) framework for PRVR. First, our Backbone-Internal Temporal Adaptation allows the last few visual transformer layers to attend over groups of neighboring frames. This provides temporally aware frame embeddings while keeping CLIP frozen and training only adaptation parameters. Second, we introduce Affinity-Weighted Gradient Propagation to address the weakly supervised nature of PRVR, softly aggregating top-$k$ frames based on text-frame affinities and propagating learning signals to multiple query-relevant frames. Our method achieves state-of-the-art performance on PRVR benchmarks, demonstrates robust cross-dataset transfer, and retrieves substantially more accurate frame-level evidence within ground-truth query-relevant moments. Our code is available at github.com/hynnsk/ITA.
发表机构
- Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。