arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReLMCodec:基于预量化音素结构设计可预测的语音 token

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

Zixiang Wan, Xusheng Yang, Zheng Wang, Peiji Yang

arXiv 2608.08286首次发表:更新:

发表机构

Peking University; Tencent(北京大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReLMCodec 基于预量化音素结构,通过保留-控制-细化原则设计低比特率单码本语音编解码器,提升了语音 token 的可预测性与重建性能,且利于下游 TTS 合成。

AI 中文摘要

在语言模型时代,神经语音编解码器面临一个根本矛盾:支持高保真重建的 token 未必易于自回归模型预测。我们对多种编解码器和自监督语音表示的控制分析表明,离散码分配前更清晰的音素结构始终与更易的自回归 token 预测相关。但仅音素结构不足以实现高保真重建,还需与重建相关的声学细节。基于此观察,我们提出 ReLMCodec,这是一种基于「保留-控制-细化」原则的低比特率单码本语音编解码器:它在量化器输入处保留冻结自监督学习(SSL)特征的语言组织,通过预量化锚点保留适配(PAPA)控制驱动重建的漂移,并仅使用训练阶段的 WavLM-Large L24 教师细化量化潜在空间以减少音素级 token 碎片化。这些组件协同工作,使声学细节支持波形重建,同时让生成的 token 序列对自回归模型可预测。在 650 和 800 bps 下,ReLMCodec 在评估中推进了经验性单流可预测性-重建前沿,其增益还可迁移至下游文本到语音(TTS)合成,提升了可懂度和说话人相似度。

英文摘要

Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows that clearer phoneme structure before discrete code assignment is consistently associated with easier autoregressive token prediction. Yet phoneme structure alone is insufficient for high-fidelity reconstruction, which also requires reconstruction-relevant acoustic detail. Guided by this observation, we introduce ReLMCodec, a low-bitrate single-codebook speech codec built upon a preserve--control--refine principle: it preserves the linguistic organization of frozen self-supervised learning (SSL) features at the quantizer input, controls reconstruction-driven drift through Pre-quantization Anchor-Preserving Adaptation (PAPA), and refines the quantized latent space with a training-only WavLM-Large L24 teacher to reduce phoneme-level token fragmentation. Together, these components allow acoustic detail to support waveform reconstruction while keeping the resulting token sequence predictable for autoregressive models. At 650 and 800 bps, ReLMCodec moves the empirical single-stream predictability--reconstruction frontier in our evaluations, with gains that carry over to downstream text-to-speech (TTS) synthesis in both intelligibility and speaker similarity.

Comments18 pages, 9 figures, 12 tables. Project page: https://github.com/ggiggit/ReLMCodec

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑