ORAV:基于多模态上下文的音视频生成基准
ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts
浏览论文内容
中文总结 AI 辅助
提出ORAV基准,含380个多参考音视频生成任务,并开发参考感知成对评估协议,以衡量生成的可控性、组合性和参考忠实度。
中文摘要 AI 辅助
利用异构多模态参考进行音视频生成已成为一项新挑战,这要求对生成过程进行组合控制,并对多模态上下文进行深入理解。在本文中,我们提出了ORAV Bench(全参考音视频生成基准),包含380个任务实例,每个实例涉及2-10个参考、9种语义角色和30种角色组合。指令规定了参考之间的关系;媒体提供了要实现的身份、动态和音频特征。为了评估这些开放式输出,我们开发了一种参考感知的成对协议,该协议准备视觉和听觉证据,比较每个参考的预期贡献,并以两种呈现顺序检查总体结论。在保留实例上,该协议与人类判断的有效一致性达到86.08%。在5个前沿系统中,总体排名掩盖了不同参考组合下的独特优势。一个反复出现的失败是,尽管与参考高度相似,却复制了非预期的源内容而非请求的结果。可复现的逐点诊断(包括质量、参考亲和度和语音)揭示了模型行为的不同维度。因此,ORAV为追踪可控、组合和参考忠实的音视频生成的进展提供了一个基准。
英文摘要
Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-10 references, 9 semantic roles, and 30 role compositions. Instructions specify the relationships among references; the media supply the identities, dynamics, and audio characteristics to be realized. To evaluate these open-ended outputs, we develop a reference-aware pairwise protocol that prepares visual and auditory evidence, compares the intended contribution of each reference, and checks the overall verdict in both presentation orders. On held-out instances, it achieves 86.08% effective agreement with human judgments. Across 5 frontier systems, overall rankings conceal distinct strengths across reference compositions. A recurring failure is to reproduce unintended source content in place of the requested result, despite closely resembling a reference. Reproducible pointwise diagnostics of quality, reference affinity, and speech reveal distinct dimensions of model behavior. ORAV thus offers a benchmark for tracking progress toward controllable, compositional, and reference-faithful audio-video generation.
发表机构
- College of AI, Tsinghua University(清华大学人工智能学院)
- Tencent Hy(腾讯Hy)
机构由 AI 辅助整理,请以论文原文为准。