TokenMapper:迈向可互操作语音标记翻译的一步
TokenMapper: A Step Toward Interoperable Speech Token Translation
浏览论文内容
中文总结 AI 辅助
TokenMapper提出方向感知的离散域标记翻译框架,实现异构语音分词器间直接映射,降低延迟并保持性能,迈向跨模型语音标记互操作。
中文摘要 AI 辅助
神经音频编解码器将语音离散化为标记序列,但由此产生的标记空间在词汇表和码本结构上存在差异,阻碍了模型间的直接通信。这一限制影响了诸如对话式语音代理和语音到语音翻译系统等多语音模型必须交互的应用。因此,在语音系统之间传递信息通常需要解码为波形音频,再用第二个分词器重新编码,这增加了延迟并引入了潜在的信息丢失。为解决这些限制,我们提出了TokenMapper,一个方向感知的框架,用于在离散域中实现异构语音分词器之间的直接标记到标记翻译。TokenMapper支持结构不匹配的标记空间,包括在共享有效标记率下,单码本和多码本表示之间的映射。在GLM-4-Voice、MiMi和DualCodec上的实验显示了一致的跨模型性能。具体来说,翻译的词错误率(WER)接近原生重建,绝对WER在2.5-6.8%以内;TokenMapper输出的人工MOS评分范围为2.29至4.39,与UTMOS的方向级别趋势一致,端到端延迟相对于波形桥接降低了4.8-94.5%,每个话语最高可达972毫秒。这些结果为无需中间波形重建的跨模型语音标记互操作性提供了切实的一步。
英文摘要
Neural audio codecs discretize speech into token sequences, but the resulting token spaces differ in vocabulary and codebook structure, preventing direct communication across models. This limitation affects applications such as conversational voice agents and speech to speech translation systems where multiple speech models must interact. As a result, transferring information between speech systems typically requires decoding to waveform audio and re-encoding with a second tokenizer, increasing latency and introducing potential information loss. To address these limitations, we present TokenMapper, a direction aware framework for direct token to token translation between heterogeneous speech tokenizers in the discrete domain. TokenMapper supports structurally mismatched token spaces, including mappings between single codebook and multi codebook representations, under a shared effective token rate. Experiments on GLM-4-Voice, MiMi and DualCodec show consistent cross model performance. Specifically, translation WER approaches native reconstructions within 2.5-6.8% absolute WER, human MOS for TokenMapper outputs ranges from 2.29 to 4.39, following the same direction level trends as UTMOS and end to end latency is reduced by 4.8-94.5% relative to waveform bridging, reaching up to 972 ms per utterance. These results provide a practical step toward cross model speech token interoperability without intermediate waveform reconstruction.
发表机构
- Ben-Gurion University of the Negev(内盖夫本-古里安大学)
机构由 AI 辅助整理,请以论文原文为准。