arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25300cs.CV

MEDit-Bench:一个用于评估消息驱动叙事视频编辑的数据集

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

发表机构大阪大学 · 株式会社CyberAgent
查看机构详情
  • The University of Osaka(大阪大学)
  • CyberAgent, Inc.(株式会社CyberAgent)

机构由 AI 辅助整理,请以论文原文为准。

Katsuya Ogata, Zongshang Pang, Mayu Otani, Yuta Nakashima

首次发表
浏览论文内容

中文总结 AI 辅助

研究消息驱动的视频编辑问题,提出MEDit-Bench数据集及基准,通过多消息与编辑配对展示消息对编辑的影响,定义评估协议,标注消息属性,实验表明模型虽在宽松阈值下接近人类,但严格标准下仍落后,人类编辑更优。

中文摘要 AI 辅助

视频编辑本质上是由消息驱动的:即使来自相同的源素材,根据编辑想要传达的叙事,所选镜头也会改变。视频摘要这一密切相关任务的基准将编辑意图简化为单一的、与消息无关的显著性概念,因此无法体现这种多样性。为评估消息驱动的视频编辑,我们提出MEDit-Bench,一个数据集和基准,它将长视频与多个编辑消息以及每个消息的多个专业制作的编辑配对,表明不同消息会从相同源产生截然不同的编辑。我们基于时间对齐指标定义了一个自动评估协议,并发现基于大语言模型的判断偏好作为叙事质量的自然代理,由于严重的位置偏差,在这项任务中不可靠。我们还为每个消息标注了模糊性和上下文性分数,并表明这两个维度与模型性能呈负相关,将消息难度确立为一个有意义的分层因素。使用最先进的多模态大语言模型和强化微调基线进行的实验表明,虽然强大的模型在宽松阈值下接近人类时间对齐,但在更严格的标准下,所有模型都落后于人类。一项人类感知研究进一步证实了巨大的质量差距,专业人类编辑始终比模型输出更受青睐。

英文摘要

Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a closely related task, video summarization, reduce editorial intent to a single, message-agnostic notion of saliency and thus do not account for this diversity. For evaluating message-driven video editing, we present \textbf{MEDit-Bench}, a dataset and benchmark, which pairs long-form videos with multiple editing messages and multiple professionally produced edits per message, demonstrating that different messages yield substantially different edits from the same source. We define an automatic evaluation protocol based on temporal alignment metrics, and find that an LLM-as-a-judge preference, a natural proxy for narrative quality, is unreliable for this task due to severe position bias. We additionally annotate each message with ambiguity and contextfulness scores, and show that both dimensions negatively correlate with model performance, establishing message difficulty as a meaningful stratification factor. Experiments with state-of-the-art MLLMs and reinforcement fine-tuned baselines show that while strong models approach human temporal alignment at lenient thresholds, all models fall behind humans at stricter criteria. A human perceptual study further confirms a large quality gap, with professional human edits remaining consistently preferred over model outputs.

↑