AI 中文总结
针对现实声学条件下西班牙语语音理解受关注少的问题,引入ESCUCHA基准测试,含1000个人工挑选与音频配对的问题,强调推理与语言多样性,涵盖多类问题,测试发现先进模型与人类存在性能差距。
AI 中文摘要
随着大型音频语言模型(LALMs)的发展,强大的评估框架变得至关重要。在此背景下,现实声学条件下的西班牙语语音理解受到的关注特别少。我们引入了ESCUCHA,这是首个旨在评估LALMs在异构声学条件和推理能力方面的西班牙语语音理解基准测试。ESCUCHA包含1000个由人工挑选的与音频配对的问题,总计162.9小时的音频直接来自“真实场景”而非现有数据集,时长从几秒到80多分钟不等。该基准测试强调推理,涵盖9个感知和10个推理类别,并通过多种西班牙口音和非规范语音体现语言多样性。ESCUCHA还包括多音频问题、口语问题和音频指令,并标记出支持开放式评估的问题。对几个先进的多模态和语音模型进行基准测试发现,它们与训练有素的人类相比存在显著的性能差距。
英文摘要
As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LALMs across heterogeneous acoustic conditions and reasoning abilities. ESCUCHA comprises 1,000 human-curated questions paired with audio, totaling 162.9 hours sourced directly ``from the wild'' rather than drawn from existing datasets, with durations ranging from a few seconds to over 80 minutes. The benchmark emphasizes reasoning, spanning 9 perceptual and 10 reasoning categories, and it captures linguistic diversity through multiple Spanish accents and non-normative speech. ESCUCHA further includes multi-audio questions, spoken questions, and audio instructions, and it flags which questions support open-ended evaluation. Benchmarking several state-of-the-art multimodal and speech models reveals substantial performance gaps relative to trained humans.
CommentsUnder review