将语音语言模型转变为多语言听众
Turning Speech Language Models into Multilingual Listeners
浏览论文内容
中文总结 AI 辅助
针对语音语言模型多语言能力不足的问题,提出大规模合成数据集MULTISPEECHQA和基准MULTISPEECH-BENCH,通过微调Qwen 2.5-Omni提升性能,证明合成数据是廉价有效的解决方案。
中文摘要 AI 辅助
语音语言模型(SLMs)能够理解口语问题,但仅支持少数几种高资源语言,限制了全球数百万人的使用。这一差距源于多语言语音指令调优数据集的稀缺。我们提出了MULTISPEECHQA,这是一个大规模、合成生成并经人工验证的数据集,包含23种类型多样的语言中9200小时、1080万个口语问答对。利用MULTISPEECHQA,我们还引入了MULTISPEECH-BENCH,一个用于评估SLM在23种语言上性能的多任务基准。我们在MULTISPEECH-BENCH上比较了级联系统与开放权重和封闭SLM的性能,发现级联系统优于开放权重SLM,但并非所有封闭SLM。我们使用MULTISPEECHQA对Qwen 2.5-Omni进行微调,提高了其在我们基准上的性能。我们的研究结果表明,高质量的合成数据集为提升SLM的多语言能力提供了一种廉价的解决方案。
英文摘要
Speech Language Models (SLMs) that understand spoken language questions support only a few high-resource languages, limiting access to millions of people worldwide. This gap stems from the scarcity of multilingual speech instruction-tuning datasets. We present MULTISPEECHQA, a large-scale, synthetically generated and human-verified dataset comprising 9200 hours of 10.8 million spoken question-answer pairs in 23 typologically diverse languages. Using MULTISPEECHQA, we also introduce MULTISPEECH-BENCH, a multi-task benchmark for evaluating SLM performance on 23 languages. We compare the performance of a cascading system to open-weight and closed SLMs on MULTISPEECH-BENCH and find that the cascading system outperforms open-weight SLMs but not all closed SLMs. We use MULTISPEECHQA to finetune Qwen 2.5-Omni, which improves its performance on our benchmark. Our findings show that high-quality synthetic datasets offer a cheap solution to improving the multilingual capabilities of SLMs.