发表机构
ETRI; KAIST; Kyungpook National University(电子通信研究所; 韩国科学技术院; 庆北国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有文本-视频检索数据集无法适配含多事件的长视频的问题,推出首个大规模多事件文本到长视频检索数据集MELON,并提出多事件感知损失,提升了检索准确率。
AI 中文摘要
现有文本-视频检索数据集主要由包含单一主导事件的短视频片段构成,这类数据集虽适合评估基础的视觉-语言对齐能力,但难以捕捉现实检索场景——现实中长视频自然包含多个语义不同的事件,单一文本查询可能对应多个不连续的时间片段。为弥合这一差距,我们推出MELON,首个旨在将文本-视频检索扩展至具有复杂多事件结构的长视频的大规模数据集。MELON为每个视频明确标注多个事件区间及其对应的文本描述,支持对未修剪长视频的多事件理解进行训练与评估。此外,我们提出一种多事件感知损失,该损失鼓励模型区分完整事件与部分事件的匹配,从而大幅提升检索准确率。MELON数据集与所提损失共同为将文本-视频检索扩展至复杂长视频场景奠定了坚实基础,并为该领域未来研究提供了更贴近现实的评估环境。
英文摘要
Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they are limited in capturing real-world retrieval scenarios, where long-form videos naturally contain multiple semantically distinct events and a single text query may correspond to several non-contiguous temporal segments. To bridge this gap, we introduce MELON, the first large-scale dataset designed to extend text-video retrieval to long-form videos featuring complex, multi-event structures. MELON explicitly annotates multiple event intervals per video along with their corresponding textual descriptions, enabling both training and evaluation of multi-event understanding in long, untrimmed videos. In addition, we propose a multi-event aware loss that encourages models to differentiate between full-event and partial-event matches, yielding substantial improvements in retrieval accuracy. Together, the MELON dataset and our proposed loss establish a robust foundation for expanding text-to-video retrieval to complex long-form scenarios and provide a more realistic evaluation setting for future research in the field.