arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniVBench:面向全参考到视频生成的基准与大规模数据集

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Wenxue Li, Peiyan Guan, Haoyang Jiang, Junxian Cai, Hualuo Liu, Chunjie Zhang, Chong Guan, Kai Huang, Songlian Li, Taiyi Wu, Yongjian Yu, Xiaotong Zhao, Alan Zhao, Eric Liu, Xi Chen, Yu Liu, Lei Zhu

arXiv 2609.22069首次发表:更新:

发表机构

Tencent; The Hong Kong University of Science and Technology (Guangzhou)(腾讯; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有基准无法全面评估全参考到视频生成的问题,本文提出OmniVBench基准和Omni-R2V大规模数据集,涵盖7个任务族、18个细粒度任务及12,172个检查项,并揭示当前模型性能差距。

AI 中文摘要

参考到视频(R2V)生成正朝着日益通用和灵活的参考控制方向发展,催生了全参考到视频(omni R2V)生成这一新兴范式。然而,现有基准不足以覆盖这些新兴能力:其测试用例涵盖的参考类型和组合有限,且评估协议主要评估整体参考一致性,忽视了参考因素是否被正确保留、解耦和路由。同时,构建全参考到视频训练数据的高昂成本使得合适的训练资源稀缺。为弥补这些不足,我们提出了OmniVBench和Omni-R2V数据集,用于评估和训练全参考到视频模型。OmniVBench将R2V评估扩展到更广泛的参考类型、细粒度控制任务和更丰富的参考组合,涵盖7个任务族和18个细粒度任务,涉及内容、运动、风格、结构、叙事和多参考设置。我们引入了基于因素的评估,包含12,172个针对具体案例的检查项,评估预期参考因素是否被忠实保留、正确解耦并绑定到其目标,以及是否根据指令正确实现。我们进一步推出了Omni-R2V数据集,为更广泛的研究社区提供多样化的R2V任务的工业级训练资源。该数据集主要基于大规模专业视频素材库,包含340K个经过处理的训练样本,涵盖多种参考类型和多参考组合。我们开发了针对具体任务的参考-目标对构建流程,为全参考到视频数据构建提供了一种实用且可扩展的方案。对先进的开源和闭源R2V模型的广泛评估揭示了在OmniVBench上各任务族和评估维度上的明显性能差距,凸显了当前R2V模型仍存在的局限性。

英文摘要

Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. We introduce factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. We develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals clear performance gaps across task families and evaluation dimensions on OmniVBench, highlighting remaining limitations of current R2V models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑