VEFX-Bench:一个全面的通用视频编辑和视觉特效基准测试
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
浏览论文内容
中文总结 AI 辅助
本文提出VEFX-Bench,基于VEFX-Dataset构建,用于标准化比较视频编辑系统,通过VEFX-Reward模型提升评估准确性,揭示当前模型在视觉合理性、指令遵循和编辑局部性方面的差距。
中文摘要 AI 辅助
随着AI辅助视频制作日益实用,指导式视频编辑成为精炼生成或捕捉素材以满足专业需求的关键。然而,该领域仍缺乏大规模人工标注的数据集和标准化评估器。本文引入VEFX-Dataset,包含5049个视频编辑示例,涵盖9大类32子类,每个示例沿三个解耦维度标注:指令遵循、渲染质量和编辑独占性。基于该数据集,提出VEFX-Reward奖励模型,联合处理源视频、编辑指令和编辑后的视频,通过序数回归预测各维度质量评分。进一步发布VEFX-Bench,包含300个精心挑选的视频提示对,用于标准化比较编辑系统。实验表明,VEFX-Reward在标准IQA/VQA指标和分组偏好评估中比通用VLM判官和先前奖励模型更符合人类判断。使用VEFX-Reward作为评估器,基准测试了代表性商业和开源视频编辑系统,揭示了当前模型在视觉合理性、指令遵循和编辑局部性方面的持续差距。项目页面为https://xiangbogaobarry.github.io/VEFX-Bench/。
英文摘要
As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large-scale human-annotated dataset with complete editing examples and a standardized evaluator for comparing editing systems. Existing resources are limited by small scale, missing edited outputs, or the absence of human quality labels, while current evaluation often relies on expensive manual inspection or generic vision-language model judges that are not specialized for editing quality. We introduce VEFX-Dataset, a human-annotated dataset containing 5,049 video editing examples across 9 major editing categories and 32 subcategories, each labeled along three decoupled dimensions: Instruction Following, Rendering Quality, and Edit Exclusivity. Building on VEFX-Dataset, we propose VEFX-Reward, a reward model designed specifically for video editing quality assessment. VEFX-Reward jointly processes the source video, the editing instruction, and the edited video, and predicts per-dimension quality scores via ordinal regression. We further release VEFX-Bench, a benchmark of 300 curated video-prompt pairs for standardized comparison of editing systems. Experiments show that VEFX-Reward aligns more strongly with human judgments than generic VLM judges and prior reward models on both standard IQA/VQA metrics and group-wise preference evaluation. Using VEFX-Reward as an evaluator, we benchmark representative commercial and open-source video editing systems, revealing a persistent gap between visual plausibility, instruction following, and edit locality in current models. Our project page is https://xiangbogaobarry.github.io/VEFX-Bench/.
发表机构
- Texas A\&M University(德克萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。