通过迭代S2U-T2U优化得到的说话人归一化语义语音令牌
NormToken: Speaker- and Duration-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement
浏览论文内容
中文总结 AI 辅助
本研究提出迭代语义令牌纯化(ISTP)方法,通过交替S2U与T2U训练优化语义语音令牌,提升了S2U-T2U一致性,在语音转换等任务中表现优异。
中文摘要 AI 辅助
语义语音令牌应保留语言内容,同时抑制从声学输入继承的说话人及时长相关的变化。我们提出迭代语义令牌纯化(ISTP)方法,这是一种由文本可预测性引导的交替语音到单元(S2U)和文本到单元(T2U)训练流程。从初始S2U令牌化器开始,每次迭代在其去重后的令牌序列上训练T2U模型,解码后的T2U预测作为新初始化S2U模型的连接时序分类目标,该S2U模型的输出又用于监督下一个T2U模型。此循环逐步对齐两个令牌生成器,使令牌空间偏向可从文本恢复的信息。在普通话和英语上的实验显示,S2U与T2U的一致性大幅提升。独立训练的反令牌化器进一步表明,优化后的S2U和T2U令牌保留了足够内容,可用于高可懂度的语音转换和文本到语音合成。在语音转换中,生成的语速更贴近参考;优化后的令牌还展现出大幅提升的跨说话人一致性,且减少了可被探针恢复的说话人信息。
英文摘要
Semantic speech tokens should preserve linguistic content while suppressing utterance-specific acoustic and duration variation. However, existing speech-to-unit (S2U) tokenizers often retain speaker-related acoustic characteristics and duration information. To address this issue, we propose NormToken, an iterative semantic token purification framework that alternates S2U and text-to-unit (T2U) training. In each iteration, the T2U model produces text-derived tokens to supervise a newly initialized S2U tokenizer. The resulting S2U tokens are then used as targets for the next T2U iteration. This cycle drives the two models toward a shared, text-predictable token space. Experiments on Mandarin and English demonstrate improved S2U--T2U agreement and parallel-utterance token consistency. De-tokenizers trained on the initial and refined tokens further show that refined tokens maintain comparable WER and CER while improving speaker similarity in both voice cloning and text-to-speech synthesis. In voice cloning, refined tokens produce speaking rates closer to the acoustic reference, suggesting reduced dependence on source duration information. Audio samples are available at https://hanlin1004.github.io/normtoken_demopage.
发表机构
- Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。