arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VideoArgus:基于智能体评分规则的视频生成与编辑统一评估框架

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

Ziyun Zeng, Zixuan Wang, Yongsheng Yu, Hang Hua, Jiebo Luo

arXiv 2608.05485首次发表:更新:

发表机构

University of Rochester(罗切斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视频评估基准的局限,本文提出VideoArgus框架,构建含1026个实例的VideoArgus-Bench,其在与人类判断的相关性上优于基准特定评估器,且模型排名稳定。

AI 中文摘要

评估生成视频仍存在挑战,因为现有基准依赖固定评估内容,仅覆盖部分生成与编辑场景,且评分证据有限。本文提出VideoArgus,一个覆盖5种视频生成与编辑场景的统一评分规则驱动框架。对于每个输入实例,VideoArgus生成一次不依赖输出、针对样本的评分规则,并复用该规则评估所有对应候选视频。该评分规则定义具体标准、评分规则、失败模式及证据计划,指导特定标准的VLM(视觉语言模型)问答与视觉工具生成基于证据的标准评分、理由及诊断报告。我们进一步构建VideoArgus-Bench,包含1026个精心整理的输入实例,由653张高质量图像和416个高质量视频构建,所有基准评分规则均已预生成、冻结并发布。在独立的1260个视频的人类对齐集上,VideoArgus在所有5项任务中,与人类判断的输入内Spearman和Kendall相关性均高于对应基准特定评估器。不同评分规则生成及评估VLM主干下,模型排名也基本保持一致。所有代码和数据均已发布,访问我们的项目页面:this https URL

英文摘要

Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output-blind, sample-specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion-specific VLM QA and visual tools to produce evidence-grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus-Bench, containing 1,026 curated input instances built from 653 high-quality images and 416 high-quality videos, with all benchmark rubrics pre-generated, frozen, and released. On a separate 1,260-video human-alignment set, VideoArgus achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric-generation and evaluation-VLM backbones. All code and data are released. Visit our project page: https://zzzmyyzeng.github.io/VideoArgus

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑