arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EviDETR:保留查询相关时间证据用于时刻检索与高亮检测

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang

arXiv 2609.30724首次发表:更新:

发表机构

Beijing Normal-Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EviDETR通过语义感知特征重加权、时间Top-2混合专家解码器和MR2HD融合保留查询相关证据,在QVHighlights等基准上显著提升时刻检索与高亮检测性能。

AI 中文摘要

联合视频时刻检索与高亮检测需要在估计片段级显著性的同时识别查询相关的时间段,然而DETR风格的流水线在编码、解码和跨任务预测过程中并未显式保留查询相关的证据。我们提出EviDETR,一个证据保留框架,包含三个组件。语义感知特征重加权(SFR)通过显著性估计和跨模态交互增强查询相关的片段表示。时间Top-2混合专家(TTop2MoE)解码器通过稀疏专家路由进行查询自适应细化。MR-to-HD(MR2HD)融合通过置信度加权的多尺度聚合将跨度级检索证据转移到片段级高亮预测。使用CLIP+SlowFast特征,EviDETR在QVHighlights上实现了时刻检索的69.29 R1@0.5、54.77 R1@0.7和48.41平均mAP,以及41.83 HD-mAP和68.33 HIT@1。在TACoS和Charades-STA上的强结果进一步证明了跨数据集的可迁移性。

英文摘要

Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.

Comments5 pages, 3 tables, 1 figure. Submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑