arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过资源感知型语音编码器混合模型打破多对多语音到文本翻译中的多语言性诅咒

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu

arXiv 2608.04586首次发表:更新:

AI 中文总结

针对多对多语音到文本翻译中的多语言性诅咒问题,提出资源感知型MoSE框架与五阶段课程学习策略,其4B参数模型在45种语言的全方向翻译任务上实现最优性能,显著提升低资源语言表现且不损害高资源性能。

AI 中文摘要

多模态大语言模型(MLLMs)在语音到文本翻译(S2TT)领域已取得显著成功。然而,当处理多语言语音输入时,跨所有语言共享的单一语音编码器会面临多语言性诅咒问题:不同资源水平的语言会争夺有限的表示能力,导致高资源语言性能强劲,但低资源语言的语音性能大幅下降。为解决这一问题并提升多语言一致性,我们提出MSRT,这是一个围绕资源感知型语音编码器混合模型(Mixture of Speech Encoders, MoSE)构建的新型框架。MoSE使用显式语言路由器将每个语音分配给合适的专家编码器:一个冻结的专家保留高资源语言能力,一个可训练的专家则适配并专攻中、低资源语言。我们进一步引入五阶段课程学习策略,大幅降低了数据依赖性,每种语言仅需10小时的配对S2TT数据即可实现有效对齐。我们在45种语言上开展了广泛实验,系统评估了所有45×44种翻译方向。我们的4B参数模型取得了最先进的性能,显著优于更大规模的基线模型。实证分析显示,MoSE同时提升了高、中、低资源语言的性能,其中低资源语言的提升幅度最大,从而在不损害高资源语言性能的前提下打破了多语言性诅咒。为支持未来的多语言S2TT研究,我们发布了代码和模型。

英文摘要

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substantial degradation on low-resource speech. To address this problem and improve multilingual consistency, we propose MSRT, a novel framework built around a resource-aware Mixture of Speech Encoders (MoSE). MoSE uses an explicit language router to assign each utterance to an appropriate expert encoder. A frozen expert preserves high-resource language capabilities, while a trainable expert adapts to and specializes in medium- and low-resource languages. We further introduce a five-stage curriculum learning strategy that substantially reduces data dependence, requiring only 10 hours of paired S2TT data per language for effective alignment. We conduct extensive experiments on 45 languages, systematically evaluating all $45 \times 44$ translation directions. Our 4B-parameter model achieves state-of-the-art performance, outperforming substantially larger baselines. Empirical analyses show that MoSE improves high-, medium-, and low-resource languages simultaneously, with the largest gains on low-resource speech, thereby breaking the curse of multilinguality without compromising high-resource performance. To support future multilingual S2TT research, we release our code and models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑