arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SEA-SpeechBench:面向东南亚语音理解的大规模多任务基准

SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw

arXiv 2609.09672首次发表:更新:

发表机构

Institute of Advanced Intelligence and Computing, A*STAR; Nanyang Technological University; Center for AI Safety(A*STAR 高级智能与计算研究所; 南洋理工大学; 人工智能安全中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对东南亚语音理解评估缺失,提出首个覆盖11种语言、9个任务的大规模多任务基准SEA-SpeechBench,并揭示现有模型在时间理解、情感识别及低资源语言上的显著性能差距。

AI 中文摘要

音频和多模态大语言模型的快速发展解锁了变革性的语音理解能力,然而评估框架仍以英语为中心,导致东南亚语言严重缺乏代表性。我们推出了SEA-SpeechBench,据我们所知,这是首个大规模多任务基准,通过99个评估集中的97,194个样本和597小时的精选音频数据,评估11种东南亚语言的语音理解能力。该基准包含3大类的9个不同任务:语音处理(自动语音识别、语音翻译、口语问答)、副语言分析(情感、性别、年龄、说话人识别)以及时间理解——一个新颖的维度,涵盖时间戳内容查询和长达3分钟的扩展音频序列中的时间定位。我们采用东南亚本土语言和英语的多语言提示,以反映用户与音频语言模型的交互方式。对领先的开源和专有系统的评估揭示了显著的性能差距。在所有模型中,时间理解、情感识别和语音翻译方面的表现仍不尽如人意。使用缅甸语和泰米尔语等低资源语言进行提示时,性能比英语落后多达41个百分点。我们的研究结果暴露了关键模型局限性,并强调了包容性模型开发的必要性。SEA-SpeechBench基准可在以下网址获取:此https URL。

英文摘要

The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA) languages critically underrepresented. We introduce SEA-SpeechBench, to the best of our knowledge, the first large-scale multitask benchmark that evaluates speech understanding in 11 SEA languages through 97,194 samples across 99 evaluation sets and 597 hours of curated audio data. Our benchmark comprises 9 diverse tasks across 3 categories: speech processing (automatic speech recognition, speech translation, spoken question answering), paralinguistic analysis (emotion, gender, age, speaker recognition), and temporal understanding, a novel dimension featuring timestamped content queries and temporal localization within extended audio sequences up to 3 minutes. We implement multilingual prompting in both native SEA languages and English to reflect user interactions with audio-language models. Evaluation of leading open-source and proprietary systems reveals marked performance gaps. Across all models, performance remains underwhelming on temporal understanding, emotion recognition, and speech translation. Prompting in low-resource languages such as Burmese and Tamil lags behind English by up to 41 percentage points. Our findings expose critical model limitations and underscore the need for inclusive model development. The SEA-SpeechBench benchmark is available at https://zwenyu.github.io/SEA-SpeechBench/.

CommentsAccepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑