发表机构
Center for Augmented Intelligence, Fondazione Bruno Kessler; Department of Information Engineering and Computer Science, University of Trento(增强智能中心,布鲁诺·凯斯勒基金会; 特伦托大学信息工程与计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MEUSLI是首个多语言投影器家族,连接Whisper编码器与开源多语言大语言模型,实现28种欧洲语言的开源端到端自动语音识别。它扩展了单语言管道,能借助持续学习技术扩展到其他语言,还可用于语音翻译和主题识别等任务。
AI 中文摘要
轻量级投影器是连接预训练语音编码器和大语言模型的既定方式,可将声学特征映射为用于自动语音识别和口语问答等任务的令牌级嵌入。然而,现有系统通常仅支持少数语言且常限于英语。我们引入MEUSLI,首个将Whisper编码器与开源多语言大语言模型相连的开放科学多语言投影器家族,实现28种欧洲语言的完全开源端到端自动语音识别。MEUSLI扩展了先前的单语言管道,在高资源和低资源语言中均有出色表现。利用适当的持续学习技术,MEUSLI可轻松扩展到训练中未见过的其他语言。我们还证明MEUSLI投影器可用于自动语音识别之外的任务,通过每种语言仅几小时的特定任务监督实现多语言语音翻译和主题识别。总体而言,MEUSLI为多语言语音理解任务提供了坚实基础,支持可扩展且包容的开源语音大语言模型。
英文摘要
Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for tasks like ASR and spoken question answering. Existing systems, however, typically only support a few languages and are often limited to English. We introduce MEUSLI, the first open-science multilingual projector family that links a Whisper encoder with open-source multilingual LLMs, enabling fully open-source end-to-end ASR in 28 European languages. MEUSLI extends prior monolingual pipelines, delivering strong results across high- and low-resource languages. Using proper continual leaning techniques, MEUSLI can be easily extended to other languages not seen in training. We further demonstrate that the MEUSLI projector can be leveraged beyond ASR, enabling multilingual speech translation and topic identification with only a few hours of task specific supervision per language. Overall, MEUSLI provides a solid foundation for multilingual speech understanding tasks, supporting scalable and inclu- sive open-source SpeechLLM