arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MMAC:一个用于音频字幕生成的大规模多维度基准测试集

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Weijie Wu, Junbo Li, Lin Li, Jun Fang, Qingyang Hong

arXiv 2607.27109首次发表:更新:

AI 中文总结

该研究针对音频字幕生成评估的信息覆盖与可靠性诊断难题,构建了含5638个音频的MMAC多维度基准,评估多种AudioLLMs并将发布基准与代码。

AI 中文摘要

随着音频大语言模型(AudioLLMs)的发展,音频字幕生成需要从简短描述转向开放式、细粒度的自由形式描述。现有评估通常聚焦于生成质量或任务性能,难以诊断信息覆盖度和描述可靠性。我们提出MMAC,即大规模多维度音频字幕生成基准测试集。MMAC包含来自20多个数据源的5638个音频片段,覆盖6个能力类别和15个评估维度。对于模型生成的字幕,MMAC会检查其是否提及目标维度的相关信息,以及提及内容是否与参考标签一致。我们评估了代表性的开源和闭源AudioLLMs,结果显示不同评估维度、信息覆盖度和描述可靠性存在明显差异。我们将发布MMAC基准测试集及评估代码。

英文摘要

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑