发表机构
Human Language Technology Center of Excellence; Johns Hopkins University; Georgetown University; George Mason University; University of Rochester(人类语言技术卓越中心; 约翰斯·霍普金斯大学; 乔治城大学; 乔治梅森大学; 罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对原始视频缺乏上下文导致检索与理解困难的问题,本文提出多语言基准MultiVENT-Raw,含近12万原始视频及事件查询,支持检索与生成任务,并证明对最新多模态模型仍有挑战。
AI 中文摘要
在线信息越来越多地以视频形式被消费。其中很大一部分以*原始视频*的形式出现:即用手机、手持摄像机或通过闭路电视拍摄的连续镜头,随后直接上传到社交媒体平台和内容共享服务。专业或甚至业余剪辑的素材通常包含脚本化语音、字幕、图形和有助于将主题情境化的元数据,而原始视频通常不包含这些元素,使其成为信息检索和机器理解更具挑战性的媒介。为促进该领域的发展,我们发布了MultiVENT-Raw,这是一个包含近12万个主要为原始视频(总时长超过5300小时)的多语言集合,配有130个事件和222个以事件为中心的查询,以及人工标注的视频相关性判断和针对相关视频的人工提取关键事实。MultiVENT-Raw支持检索任务(识别集合中与查询事件相关的视频)和生成任务(将事件相关视频总结成面向目标用户的连贯报告)。我们在MultiVENT-Raw上对强基线进行了基准测试,表明即使对于一些最新的多模态模型,这两个任务也具有挑战性。
英文摘要
Online information is increasingly consumed in video format. Much of this comes in the form of *raw video*: continuous footage taken on a cell phone, with a hand-held camera, or via CCTV, which is then directly uploaded to social media platforms and content sharing services. Whereas professional or even amateur-edited footage tends to feature scripted speech, chyrons, graphics, and metadata that help contextualize its subject matter, raw video typically contains none of these things, making it a much more challenging medium for information retrieval and machine understanding. To facilitate progress in this domain, we release MultiVENT-Raw, a multilingual collection of nearly 120,000 primarily raw videos (over 5,300 total hours), paired with 130 events and 222 event-centric queries, along with human-annotated video relevance judgments and human-extracted key facts for relevant videos. MultiVENT-Raw supports both a retrieval task---to identify videos in the collection relevant to a query event---and a generation task---to summarize event-related videos into a coherent report for a target user. We benchmark strong baselines on MultiVENT-Raw, showing both tasks to be challenging even for some of the latest multimodal models.