MuLA-Bench:通过多层审计构建的多语言长格式音频理解基准
MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing
浏览论文内容
中文总结 AI 辅助
MuLA-Bench通过多层审计构建多语言长音频理解基准,覆盖16种语言和8个领域,评估十个模型,揭示语言、证据和任务对难度的联合影响及条件性失败模式。
中文摘要 AI 辅助
长格式音频的性能通常通过上下文长度和总体准确率来概括,这掩盖了语言、证据和任务如何共同影响难度。我们引入了MuLA-Bench:包含5,038个开放式问题,覆盖1,769个野外录音,总计1,377.9小时,涵盖16种语言和8个领域。一个平衡的语言×领域语义轨道支持受控比较,而一个互补的声学轨道保留了自然发生的非语音证据。基于证据的生成、捷径检查和语言专家审查提供了可审计的问题,无需翻译共享源集或注入目标声音。我们评估了十个音频语言模型,并在一个固定的八模型队列上进行了汇总诊断。语言排名在不同领域和任务间有所变化;声学-语义性能差距随请求的操作而变化;在正确事件被识别后,时间错误可能仍然存在。长距离检索相对较强,而精确的时钟对齐和自然声学事件的事实依据仍然脆弱。因此,MuLA-Bench揭示了单一长上下文分数无法捕捉的条件性失败模式。
英文摘要
Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench: 5,038 open-ended questions over 1,769 in-the-wild recordings totaling 1,377.9 hours, covering 16 languages and eight domains. A balanced Language x Domain semantic track supports controlled comparisons, while a complementary acoustic track preserves naturally occurring non-speech evidence. Evidence-grounded generation, shortcut checks, and language-expert review provide auditable questions without translating a shared source set or injecting target sounds. We evaluate ten audio-language models and conduct pooled diagnostics on a fixed eight-model cohort. Language rankings change across domains and tasks; acoustic-semantic performance gaps vary with the requested operation; and temporal errors can persist after the correct event is identified. Long-range retrieval is comparatively strong, while precise clock alignment and factual grounding of natural acoustic events remain fragile. MuLA-Bench thus exposes conditional failure patterns that a single long-context score does not capture.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Alibaba Token Hub, Alibaba Group(阿里巴巴集团通义实验室)
机构由 AI 辅助整理,请以论文原文为准。