发表机构
The University of Hong Kong; Alibaba Group; Zhejiang University; Peking University(香港大学; 阿里巴巴集团; 浙江大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有基于指令的视频编辑基准的任务覆盖有限、指标不足的问题,OmniEdit-Bench将编辑任务分解为多维度并提出含惩罚机制的评估框架,实验显示当前IVE模型仍待提升。
AI 中文摘要
基于指令的视频编辑(IVE)是一个具有广泛应用的新兴领域,但评估编辑模型仍存在挑战。现有基准存在两个主要局限:一是继承自图像编辑的任务覆盖范围有限,忽略了视频特有的维度;二是指标不足,无法衡量指令保真度,导致因原始视频的强视觉先验,错误编辑也能获得高分。为解决这些问题,我们推出了一个综合且结构化的IVE基准。该基准将编辑任务分解为多个视频特有的维度,包括空间、时间、音频和基于参考的编辑,扩展了传统的帧级评估;还区分了显式和隐式指令,并融入基于推理的场景,以更好地反映现实需求。此外,我们提出了一个评估框架,从准确性、保留度、真实感和一致性四个互补维度评估编辑质量,结合人工评估和最先进的视觉-语言模型。为强调指令保真度,我们引入了一种感知准确性的惩罚机制,该机制将其他分数基于准确性进行调整,防止视觉上合理但错误的编辑获得过高评价。对代表性开源和商业模型的大量实验表明,当前IVE模型仍远未达到令人满意的水平。OmniEdit-Bench为评估基于指令的视频编辑提供了一个全面且可靠的测试平台,并为未来研究方向提供了见解。
英文摘要
Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major limitations: limited task coverage inherited from image editing, which overlooks video-specific dimensions, and inadequate metrics that fail to measure instruction fidelity, allowing incorrect edits to receive high scores due to strong visual priors from the original video. To address these issues, we introduce a comprehensive and structured benchmark for IVE. Our benchmark decomposes editing tasks into multiple video-specific dimensions, including spatial, temporal, audio, and reference-based editing, extending beyond conventional frame-level evaluation. It also distinguishes explicit and implicit instructions and incorporates reasoning-based scenarios to better reflect real-world requirements. Furthermore, we propose an evaluation framework that assesses editing quality from four complementary dimensions: accuracy, preservation, realism, and consistency, using both human judgments and state-of-the-art vision-language models. To emphasize instruction fidelity, we introduce an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving inflated evaluations. Extensive experiments on representative open-source and commercial models show that current IVE models remain far from satisfactory. OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions. The project page is https://omniedit-bench.github.io/.