arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MELON:用于多事件文本到长视频检索的大规模数据集

MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval

Chan Hur, SeungWoo Song, Jeong-hun Hong, Won Jun Oh, Hyeyoung Park, KyungTae Lim

arXiv 2609.01654首次发表:更新:

发表机构

ETRI; KAIST; Kyungpook National University(电子通信研究所; 韩国科学技术院; 庆北国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有文本-视频检索数据集无法适配含多事件的长视频的问题,推出首个大规模多事件文本到长视频检索数据集MELON,并提出多事件感知损失,提升了检索准确率。

AI 中文摘要

现有文本-视频检索数据集主要由包含单一主导事件的短视频片段构成,这类数据集虽适合评估基础的视觉-语言对齐能力,但难以捕捉现实检索场景——现实中长视频自然包含多个语义不同的事件,单一文本查询可能对应多个不连续的时间片段。为弥合这一差距,我们推出MELON,首个旨在将文本-视频检索扩展至具有复杂多事件结构的长视频的大规模数据集。MELON为每个视频明确标注多个事件区间及其对应的文本描述,支持对未修剪长视频的多事件理解进行训练与评估。此外,我们提出一种多事件感知损失,该损失鼓励模型区分完整事件与部分事件的匹配,从而大幅提升检索准确率。MELON数据集与所提损失共同为将文本-视频检索扩展至复杂长视频场景奠定了坚实基础,并为该领域未来研究提供了更贴近现实的评估环境。

英文摘要

Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they are limited in capturing real-world retrieval scenarios, where long-form videos naturally contain multiple semantically distinct events and a single text query may correspond to several non-contiguous temporal segments. To bridge this gap, we introduce MELON, the first large-scale dataset designed to extend text-video retrieval to long-form videos featuring complex, multi-event structures. MELON explicitly annotates multiple event intervals per video along with their corresponding textual descriptions, enabling both training and evaluation of multi-event understanding in long, untrimmed videos. In addition, we propose a multi-event aware loss that encourages models to differentiate between full-event and partial-event matches, yielding substantial improvements in retrieval accuracy. Together, the MELON dataset and our proposed loss establish a robust foundation for expanding text-to-video retrieval to complex long-form scenarios and provide a more realistic evaluation setting for future research in the field.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑