发表机构
Elm Company(埃尔姆公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Nuha-Speech构建了含150万样本的阿拉伯语语音问答语料库,基于Qwen-Omni微调模型,并设计评估框架,以建立通用阿拉伯语语音大语言模型的基础设施。
AI 中文摘要
随着语音大语言模型(speech-LLMs)日益多语言化,阿拉伯语仍然明显代表性不足,这凸显了建立专门基础设施以训练和评估阿拉伯语语音大语言模型的必要性。为解决这一差距,我们推出了Nuha-Speech,这是一项旨在开发通用阿拉伯语语音大语言模型的综合性计划,涵盖数据集构建、模型训练和系统评估。具体而言,我们构建了一个大规模阿拉伯语语音问答(SQA)语料库,包含超过150万个训练样本,以便在广泛的核心里语音任务上进行指令微调。随后,该语料库被用于基于不同规模的Qwen-Omni模型变体进行监督微调。最后,我们设计了一个评估框架,包含多样化的任务和定制化指标。通过这项工作,我们旨在在有限的阿拉伯语语音资源约束下,为阿拉伯语语音大语言模型建立基础性基础设施。
英文摘要
As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap, we introduce Nuha-Speech, a comprehensive initiative to develop general-purpose Arabic speech-LLMs spanning dataset construction, model training, and systematic evaluation. Specifically, we constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks. Then, the corpus was used for supervised fine-tuning based on Qwen-Omni model variants at different scales. Finally, we designed an evaluation framework featuring diverse tasks and tailored metrics. Through this work, we aim to establish foundational infrastructures for Arabic Speech-LLMs under constraints imposed by limited Arabic speech resources.