AI 中文总结
研究针对多参考到音频-视频生成设置,引入MultiRef-Compass基准。通过可扩展管道构建350个样本,定义含四个维度14个子指标的评估协议,集成自动指标与MLLM评判框架,实验表明该基准对MR2AV研究有重要意义。
AI 中文摘要
多参考到音频-视频(MR2AV)生成旨在根据多个参考和文本指令生成连贯的音频-视频内容。现有基准主要集中在文本驱动生成、单参考主体保留或孤立的音频-视频对齐,新兴的MR2AV设置大多未被探索。与这些设置相比,MR2AV要求模型在生成同步视觉和音频内容时联合推理多个参考。模型不仅要忠实地保留每个参考,还要正确地绑定和组合多个参考实体成连贯的视听事件。为填补这一空白,我们引入了MultiRef-Compass,一个用于MR2AV生成的统一基准。它由通过可扩展且可控的资产组合管道精心策划的350个样本组成,涵盖多视图主体保留、多实体绑定和人-对象-场景组合。为提供可解释的评估,MultiRef-Compass定义了一个具有四个维度的评估协议:基本质量、参考一致性、视听一致性和指令遵循,使用14个子指标。MultiRef-Compass将自动指标与重新评判增强的MLLM作为评判框架集成,实现对感知保真度和参考条件组合的可扩展和可审计评估。对八个代表性MR2AV系统的广泛实验揭示了多个评估维度上有很大的改进空间,强调了全面基准的必要性,并将MultiRef-Compass定位为未来MR2AV研究的基础。
英文摘要
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises $350$ carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.
Comments32 pages