arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PRISM-Bench:面向文本生成音视频的以音频为核心的诊断基准

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

Yuchen Sun, Qian Yang, Jun Wang, Detai Xin, Guoqiao Yu, Guanglu Wan, Qi Jia

arXiv 2609.04867首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory; Meituan(上海人工智能实验室; 美团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PRISM-Bench作为首个以音频为核心的T2AV诊断基准,通过多维度评估揭示了前沿与开源T2AV模型的性能差距,指出当前模型在复杂音频关联控制任务上的不足。

AI 中文摘要

文本生成音视频(T2AV)技术发展迅速,但其评估仍低估了音频模态的重要性。现有基准要么将音频视为视频质量的辅助部分,要么与音视频关联度分开评估,导致难以诊断当前系统在音频生成方面的真正成败之处。我们提出PRISM-Bench,这是首个面向T2AV生成的以音频为核心的诊断基准。该基准基于精心筛选的900个人工验证样本构建,沿两个正交轴对音频评估进行分解:音频类型(语音、音乐和音效)以及声源可见性(屏幕内与屏幕外)。它通过35项细粒度标准在四个感知维度上评估生成内容:音视频一致性、音频质量、音频表现力和提示遵循度。为确保评估可靠,我们采用了增强型的多模态大语言模型作为评判者(MLLM-as-a-Judge)协议,该协议基于与真实参考的盲法成对比较,与人类评分者的一致性达到70%以上。我们对近期T2AV系统的评估显示,前沿模型与开源模型之间存在显著的性能差距。此外,我们发现当前生成范式过度拟合感知保真度,却在复杂的关联和控制任务上表现不佳,尤其是在生成音乐和同步的屏幕内音频时。

英文摘要

Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in isolation from audiovisual grounding, making it difficult to diagnose where current systems truly succeed or fail in audio generation. We present PRISM-Bench, the first audio-centric diagnostic benchmark for T2AV generation. Built from a rigorously curated dataset of 900 human-verified samples, PRISM-Bench factorizes audio evaluation along two orthogonal axes: audio type (Speech, Music, and Sound) and sound-source visibility (On-screen vs. Off-screen). It evaluates generated content across four perceptual dimensions (Audio-Visual Coherence, Audio Quality, Audio Expressiveness, and Prompt Following) with 35 fine-grained criteria. To ensure reliable assessment, we adopt an enhanced MLLM-as-a-Judge protocol based on blind, side-by-side comparison against ground-truth references, demonstrating strong alignment (over 70% mean agreement) with human raters. Our evaluation of recent T2AV systems highlights a significant performance gap between frontier and open-source models. Furthermore, we demonstrate that current generation paradigms overfit to perceptual fidelity while struggling with complex grounding and control tasks, particularly in generating music and synchronized On-screen audio.

Comments19 pages, 10 figures, 4 tables. Accepted at ACM Multimedia 2026 (MM '26). This arXiv version includes supplementary appendices not included in the conference proceedings version

DOI:10.1145/3767308.3836100

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑