发表机构
Wuhan University of Technology; NEC Laboratories Asia Pacific; NEC Corporation; The Hong Kong Polytechnic University(武汉理工大学; NEC亚太实验室; NEC公司; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出单塔低比特率语音编解码器BiMTokenizer,结合双向状态空间主干与RSLQ,在更少参数下实现了优于双塔基线的声学重建与语义保留性能。
AI 中文摘要
语音编解码器是连续语音信号与大语言模型之间的桥梁,但面临声学保真度与语义保留之间的固有冲突。为缓解这一冲突,近期研究越来越多地采用双塔架构,通过独立编码器将语义与声学建模解耦。然而,这类双塔设计会产生大量架构开销。为避免这种复杂性,我们重新审视单塔范式,提出BiMTokenizer——一种低比特率语音编解码器(约1.1 kbps),结合双向状态空间主干与残差球形李量化(RSLQ)。双向主干强化时间建模,而RSLQ提供固定且分离良好的格状瓶颈,实现稳健的语义与声学分词,且不会出现学习码本崩溃。实验表明,在干净和噪声环境下,BiMTokenizer在低比特率编解码器基线中实现了更优的声学重建和最低的词错误率(WER),同时参数数量不到近期双塔基线的一半。此外,其稳健的语义表示在下游语音理解任务上表现出色,证实精心设计的单塔编解码器可在低比特率下保持语义-声学平衡。代码和模型权重可在此https URL获取。
英文摘要
Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, recent works increasingly adopt dual-tower architectures to decouple semantic and acoustic modeling with separate encoders. However, these dual-tower designs incur substantial architectural overhead. To avoid such complexity, we revisit the single-tower paradigm and propose BiMTokenizer, a low-bitrate speech codec (around 1.1 kbps) combining a bidirectional state-space backbone with Residual Spherical Leech Quantization (RSLQ). The bidirectional backbone strengthens temporal modeling, while RSLQ offers a fixed, well-separated lattice bottleneck for robust semantic and acoustic tokenization without learned-codebook collapse. Experiments show that BiMTokenizer achieves superior acoustic reconstruction and the lowest WER among low-bitrate codec baselines across both clean and noisy environments, while using less than half the parameters of recent dual-tower baselines. Furthermore, its robust semantic representations yield strong performance on downstream speech understanding tasks, confirming that a well-designed single-tower codec can preserve the semantic-acoustic balance at low bitrates. The code and model weights are available at https://github.com/ZhangXinWhut/BiMTokenizer.
CommentsAccepted to EMNLP 2026 Main Conference; 19 pages, 5 figures