arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缓解零样本视频时刻检索中的模态和语言风格差距

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

Jihyun Lee, Cheol-Ho Cho, Woojin Jun, Woojin Jeong, Jae-Pil Heo

arXiv 2607.19027首次发表:更新:

发表机构

Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究零样本视频时刻检索中模态和语言风格差距问题,提出基于自相似性的时刻提议和评分方法及查询感知的多模态大语言模型推理阶段,有效缓解差距,在相关基准测试中达最优性能。

AI 中文摘要

零样本视频时刻检索旨在克服传统方法的局限性,传统方法需要大规模带有文本及其相关时间跨度标注的数据集。尽管预训练视觉语言模型和多模态大语言模型取得了进展,但现有零样本视频时刻检索方法仍严重依赖查询与视频内容的相似性,易受模态和语言风格差距影响。这些差距导致不可靠的跨度提议和不稳定的时刻检索结果。为解决此问题,我们提出基于自相似性的时刻提议和评分方法,利用视频内部的内在关系实现稳健的跨度生成和评分。通过仅从视频内容推导自相似性,避免查询帧或查询字幕相似性中的噪声和不匹配模式,从而缓解模态和语言风格差距。此外,我们引入基于查询感知的多模态大语言模型推理阶段,进一步加强文本与视频的对齐。大量实验表明,Self-SiMS在零样本视频时刻检索基准测试中达到了当前最优性能。

英文摘要

Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. Despite advances in pre-trained vision-language models and multimodal large language models, existing ZMR methods still heavily depend on query-to-video content similarity, making them vulnerable to modality and language-style gaps. These gaps lead to unreliable span proposals and unstable moment retrieval results. To address this issue, we propose Self-Similarity-based Moment Proposal and Scoring that instead exploits intrinsic relationships within videos, enabling robust span generation and scoring. By deriving self-similarity only from the video content, we circumvent the noisy and mismatched patterns of query-frame or query-caption similarities, thereby mitigating both modality and language-style gaps. Furthermore, we introduce a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video. Extensive experiments demonstrate that Self-SiMS achieves state-of-the-art performance across ZMR benchmarks.

CommentsECCV 2026 (* These authors contributed equally.)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑