发表机构
Useful Sensors(有用传感器公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对轻量级孟加拉语语音识别模型失败问题,提出用BanglaBERT WordPiece词汇替换解码器词汇并调整矩阵大小的方法,经实验改进后模型在数据集上有竞争力,提供了紧凑ASR模型跨脚本适配的可扩展蓝图。
AI 中文摘要
轻量级语音识别模型对边缘部署至关重要,但像Moonshine这样高度优化的架构在孟加拉语等形态丰富的非拉丁语上常失败。本研究将其原因归结为以英语为中心的字节级分词器,它将孟加拉语单词拆分为高生育率字节链并在推理时引发自回归崩溃。为此提出新颖的词汇移植管道,用孟加拉语脚本的BanglaBERT WordPiece词汇替换解码器词汇并调整相应令牌嵌入矩阵大小。实验结果表明令牌生育率从9.16降至1.30,自回归序列长度减少85.8%,完全缓解了解码不稳定性。在882小时的Lipi - Ghor数据集上评估时,改进后的架构实现了有竞争力的21.54%的字错误率(WER)和0.0053的实时因子(RTF)。最终,本研究为紧凑ASR模型的跨脚本适配提供了可扩展、可重现的蓝图,无需资源密集型预训练。
英文摘要
Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model's English-centric byte-level tokenizer, which fragments Bengali words into high-fertility byte chains and triggers catastrophic autoregressive collapse during inference. To resolve this, a novel vocabulary transplantation pipeline is proposed to replace the decoder vocabulary with the native-script BanglaBERT WordPiece vocabulary and resize the corresponding token embedding matrix. Experimental results demonstrate a reduction in token fertility from 9.16 to 1.30. By decreasing autoregressive sequence length by 85.8%, decoding instability is entirely mitigated. When evaluated on the 882-hour Lipi-Ghor dataset, the modified architecture achieves a competitive 21.54% Word Error Rate (WER) and a Real-Time Factor (RTF) of 0.0053. Ultimately, this research provides a scalable, reproducible blueprint for cross-script adaptation of compact ASR models without the need for resource-intensive pre-training.
Comments5 pages, 2 figures. Accepted as a poster at the MusIML Workshop, ICML 2026