发表机构
Shanghai Jiao Tong University; Shanghai AI Lab(上海交通大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有音频生成基准的不足,推出MMAG混合音频生成基准,构建含4000个标注音频的数据集,提出系统评估协议,测试发现各类模型在多控制能力间存在性能权衡,为该领域提供综合基准。
AI 中文摘要
近期音频生成系统已从单模态合成发展到生成包含语音、音乐和音效的复杂声学场景,因此评估这些模型需要评估语义保真度、说话人一致性和时序控制等多种交互能力,但现有基准聚焦于孤立领域或粗粒度描述。为填补这一空白,我们推出多控制混合音频生成(MMAG)基准。MMAG包含约4000个经人工验证的音频片段,带有覆盖语音内容、说话人身份、音乐属性、声音事件及时序关系的丰富标注,还包含用于语音克隆和时间戳条件生成的专用子集。我们进一步提出系统评估协议,测量声学保真度、语音质量、语义对齐及时序准确性。对代表性智能体编排器、统一视听生成模型和原生混合音频生成器的基准测试显示,这些能力间存在显著性能权衡,没有现有模型能始终表现良好。我们的结果凸显了可控混合音频生成的剩余挑战,并确立MMAG作为未来研究的综合基准。
英文摘要
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluating these models requires assessing multiple interacting capabilities, including semantic fidelity, speaker consistency, and temporal control, yet existing benchmarks focus on isolated domains or coarse-grained descriptions. To address this gap, we introduce the Multi-control Mixed Audio Generation (MMAG) benchmark. MMAG contains approximately 4,000 manually verified audio clips with rich annotations covering speech content, speaker identity, music attributes, sound events, and temporal relationships, together with dedicated subsets for voice cloning and timestamp-conditioned generation. We further propose a systematic evaluation protocol that measures acoustic fidelity, speech quality, semantic alignment, and temporal accuracy. Benchmarking representative agentic orchestrators, unified audio-visual generation models, and native mixed-audio generators reveals substantial performance trade-offs across these capabilities, with no existing model performing consistently well. Our results highlight the remaining challenges of controllable mixed audio generation and establish MMAG as a comprehensive benchmark for future research.
Comments15 pages, 6 figures