arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BuzzASR:一个由100多个单语语音识别模型组成的群体

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

Shivam Singh, Aditya Yadavalli, Catherine Arnett, Alex Warstadt

arXiv 2609.09554首次发表:更新:

发表机构

UC San Diego; EleutherAI(加州大学圣迭戈分校; EleutherAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BuzzASR通过大规模单语微调Whisper模型,覆盖102种语言,在多数语言上降低字符错误率,并提升分词压缩率,优于Whisper-large-v3。

AI 中文摘要

我们推出了BuzzASR,这是一个针对102种语言进行自动语音识别(ASR)适配的语言特化微调Whisper模型集合。大型端到端基于Transformer的ASR模型(如Whisper)已经彻底改变了ASR领域,但大多数主流模型都是高度多语言的。因此,这些模型在训练集中代表性不足的语言上往往表现不佳。虽然长期以来人们已知通过单语数据上的简单微调可以实现有效的语言适配,但这一策略仅被应用于少数语言。我们将这一简单方法大规模扩展到FLEURS数据集覆盖的102种语言,同时实现了一种更复杂的语言适配策略,该策略整合了单语分词器替换和使用仅文本微调的数据增强。BuzzASR模型在102种语言中的77种上优于Whisper-large-v3,平均字符错误率(CER)降低了超过2.8倍。在FLEURS和Common Voice组合测试集上,我们的模型在102种语言中的27种上达到了开源系统中的最优CER。我们的分词器替换策略相比Whisper的多语言BPE,在压缩率(每token字符数)上平均提升了3.3倍,最高提升达21.7倍。我们发布了所有模型、代码和详细结果:此https URL。

英文摘要

We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often perform poorly on languages less well-represented in their training set. While it has long been known that effective language adaptation can be achieved through simple fine-tuning on monolingual data, this strategy has only been applied to a small number of languages. We massively scale up this simple approach to 102 languages covered in the FLEURS dataset, while also implementing a more complex language adaptation strategy that integrates monolingual tokenizer replacement and data augmentation using text-only fine-tuning. BuzzASR models outperform Whisper-large-v3 on 77 out of 102 languages, reducing character error rates (CER) by a factor of over 2.8 on average. Our models achieve state-of-the-art CER among open-source systems on 27 of 102 languages on the combined FLEURS and Common Voice test set. Our tokenizer replacement strategy yields an average 3.3x improvement in compression rate (characters per token) over Whisper's multilingual BPE, with gains of up to 21.7x. We release all models, code, and detailed results: https://lemn-lab.github.io/buzz-asr

CommentsAccepted at EMNLP 2026. Models: https://huggingface.co/BuzzASR ; Project page: https://lemn-lab.github.io/buzz-asr

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑