发表机构
Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TF-PRVR提出首个免训练的部分相关视频检索框架,利用冻结的视觉-语言特征构建层次表示,通过多尺度图传播查询相关性,无需任务特定训练即可在多样数据集上保持稳定性能。
AI 中文摘要
部分相关视频检索(PRVR)旨在检索包含与给定文本查询相关时刻的未剪辑视频。尽管近期取得了进展,现有的PRVR方法存在两个关键局限:固定的视频分解方案导致语义稀释,以及任务特定训练引起的源域过拟合。在本文中,我们提出了TF-PRVR,这是首个用于PRVR的免训练框架。TF-PRVR利用冻结的视觉-语言特征来构建视频特定的层次表示。它从帧级特征中提取时间语义信号,并应用基于频率的多尺度分析来识别自适应的时间边界,生成具有连贯事件级语义的层次片段。基于这些片段,TF-PRVR构建了一个统一的多尺度图,并在时间和语义相关的节点之间传播查询相关性。随后,一种时刻感知的评分策略跨尺度聚合时间对齐的相关性,强调一致支持的时刻,同时抑制孤立的虚假响应。无需任务特定训练,TF-PRVR保留了预训练视觉-语言模型的通用对齐能力,并避免了数据集特定的过拟合。大量实验表明,TF-PRVR在具有不同视觉和时间特征的数据集上表现一致,为免训练的PRVR指明了一个实用方向。
英文摘要
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and source-domain overfitting induced by task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments demonstrate consistent performance across datasets with diverse visual and temporal characteristics, suggesting a practical direction for training-free PRVR.