TAG-Bench:面向大型音频语言模型的时序音频定位基准测试
TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
浏览论文内容
中文总结 AI 辅助
本文提出TAG-Bench基准测试评估大型音频语言模型的时序音频定位能力,发现现有模型定位精度、出现次数枚举及输出格式可靠性均存在不足,将发布相关数据与代码支持后续研究。
中文摘要 AI 辅助
大型音频语言模型(LALMs)能够描述所听到的内容,但针对其定位被查询内容发生时间的能力,尚未得到系统评估。本文提出TAG-Bench,这是一个用于时序音频定位的基准测试,模型需返回所有与自然语言查询匹配的时间区间。TAG-Bench包含1750个人工验证的查询-录音对,总时长149.5小时,涵盖8个与来源相关的子集,涉及查询类别和7秒至20分钟的音频时长;22.1%的查询存在多个真实区间。在评估的21个系统中,性能最佳的模型达到31.2的mIoU,且是仅有的在两个长子集上mIoU超过20的系统,但即便是该最优模型在IoU≥0.7时的召回率仅为21.5%。此外,21个系统中有9个mIoU低于5,且所有模型在一对多查询中均存在出现次数报告不足的问题,无一超过13.2%的计数准确率。由于响应为自由格式,本文报告了解析失败率和MAE覆盖率:解析失败会作为空预测保留在mIoU、召回率、gIoU和计数指标中,但不会计入MAE。该基准测试的结果区分了精确定位、出现次数枚举和输出格式可靠性,其跨子集比较具有描述性,而非针对查询抽象或时长的受控估计。我们将发布TAG-Bench数据和评估代码以支持未来研究。
英文摘要
Large audio language models (LALMs) can describe what is heard, but their ability to localize when queried content occurs remains less systematically evaluated. We present TAG-Bench, a benchmark for temporal audio grounding in which a model returns every time interval that matches a natural-language query. TAG-Bench contains 1,750 human-verified query-recording pairs covering 149.5 hours, with eight source-dependent subsets spanning query categories and audio durations from 7 s to 20 min; 22.1% of the queries have multiple ground-truth intervals. Across 21 evaluated systems, the best-performing model achieves 31.2 mIoU and is the only system above 20 mIoU on the two long subsets, yet even this top performer reaches only 21.5% recall at IoU >= 0.7. Moreover, 9 of 21 systems fall below 5 mIoU, and every model under-reports the number of occurrences on one-to-many queries, with none exceeding 13.2% count accuracy. Because responses are free-form, we report parsing-failure rate and MAE coverage: parsing failures remain in mIoU, Recall, gIoU, and count metrics as empty predictions but do not enter MAE. The results separate precise localization, occurrence enumeration, and output-format reliability within a benchmark whose cross-subset comparisons are descriptive rather than controlled estimates of query abstraction or duration. We will release the TAG-Bench data and evaluation code to support future research.