发表机构
Université Paris Cité; ALTAE; SCIAM(巴黎西岱大学; ALTAE机构; SCIAM机构)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对Moshi语言模型的神经编解码器Mimi的2048个token语义码本,发现其ABX实验无法捕捉语义token到音素实现的映射,经与TIMIT语料对齐后证实该码本token ID可映射至多粒度音素实现,为元语言表示研究提供了关键发现。
AI 中文摘要
本文聚焦于Moshi语言模型的神经编解码器Mimi语义码本所用的2048个token词典,发现采用Mimi开展的ABX实验无法捕捉语义token到音素实现的映射。通过将Mimi表示与TIMIT语料库转录本重新对齐,本文证实该语义码本的2048个token ID可映射到四音素、三音素、二音素、音素及亚音素实现。
英文摘要
In this paper, we focus on the dictionary of 2048 tokens used in Mimi semantic token codebook, the neural codec of the Moshi language model. We show that the ABX experiment carried out with Mimi fails to capture the mapping of the semantic tokens to phone realisations. By realigning Mimi representations to the TIMIT corpus transcriptions, we show that the 2048 tokens IDs of the semantic codebook map to quadphone, triphone, biphone, phone and subphone realisations.
Comments11 pages, accepted for the Proceedings of the Third Workshop on the Bridges and Gaps between Formal and Computational Linguistics (BriGap-3), Paris 2026