想象之后集中注意力:部分相关视频检索的文本条件证据锚定
Concentrate After Imagination: Text-Conditioned Evidence Grounding for Partially Relevant Video Retrieval
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- University of Electronic Science and Technology of China(电子科技大学)
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对部分相关视频检索中查询无关的集中瓶颈,提出分数级证据验证算子TRACE,通过激活查询相关寄存器并校准局部分数,在三个基准上取得最优并提升骨干网络性能。
AI中文摘要:
部分相关视频检索(PRVR)旨在当查询仅描述短暂时刻时检索未修剪的视频。尽管近期方法改进了局部表示、不确定性建模和全局上下文,但最终排序往往仍信任最强的局部响应;因此,一个偶然相似的片段可能产生无依据的峰值。我们将此失败识别为查询无关的集中瓶颈,并提出TRACE,一种用于PRVR的分数级证据验证算子。给定查询和全局视频寄存器,TRACE激活与查询相关的寄存器,将其支持路由到帧级证据,并在局部时间选择之前平滑地边缘化替代的查询-寄存器-帧路径。与表示级特征融合不同,TRACE仅将此证据用作原始局部分数的查询条件残差校准。在ActivityNet Captions、Charades-STA和TVR上,TRACE在所有三个基准上取得了最佳SumR,并分别将DreamPRVR骨干网络提升了1.2、1.1和1.5个百分点。消融、路由损坏、难负样本和跨骨干迁移分析支持了这样的解释:增益源于查询条件证据验证,而非通用分数偏移。
英文摘要:
Partially Relevant Video Retrieval (PRVR) retrieves untrimmed videos when queries describe only short moments. Although recent methods improve local representations, uncertainty modeling, and global context, final ranking often still trusts the strongest local response; a coincidentally similar fragment can therefore produce an unsupported peak. We identify this failure as the query-agnostic concentration bottleneck and propose TRACE, a score-level evidence verification operator for PRVR. Given a query and global video registers, TRACE activates query-relevant registers, routes their support to frame-level evidence, and smoothly marginalizes alternative query-to-register-to-frame paths before localized temporal selection. Unlike representation-level feature fusion, TRACE uses this evidence only as a query-conditioned residual calibration of the original local score. On ActivityNet Captions, Charades-STA, and TVR, TRACE achieves the best SumR on all three benchmarks and improves the DreamPRVR backbone by 1.2, 1.1, and 1.5 points, respectively. Ablation, routing-corruption, hard-negative, and cross-backbone transfer analyses support the interpretation that the gains arise from query-conditioned evidence verification rather than a generic score offset.