自蒸馏发音与口音控制用于神经文本到语音
Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech
浏览论文内容
中文总结 AI 辅助
提出自蒸馏方法,用冻结主干自身输出作为教师,训练带口音生僻词阅读,无需词典或外部编辑,在多个TTS模型上显著提升口音准确率且不损害自然度。
中文摘要 AI 辅助
读取原始文本的文本到语音系统没有词典:生僻词按猜测读出。现有补救方法是在录音语音上训练阅读与口音通道,或从示例中逐词编辑。我们两者都不做。冻结的主干网络读取包含一个已能正确说出的常见词的句子,其自身输出随后作为同一句子的教师,但该词被替换为带标签的、带口音的读音;这一训练对就是全部思路。在Sarashina2.2-TTS上,筛选的评分者在Fleiss' kappa = 0.85下,对0.89的未见词听到规定口音,而假名(无法表达口音)仅为0.57;假名未赢得任何配对;自然度未受可测量损害。未经调优迁移到自回归、扩散和编码器-解码器主干,阅读能力提升,在319个词上比无编辑高0.25至0.47;在CosyVoice 2上,三分之二的词上口音正确,但在Irodori上不行;论文定位了原因。
英文摘要
Text-to-speech that reads raw text has no lexicon: a rare word is read as guessed, and a native Japanese listener accepts a word only if its reading and pitch accent are both right. A known remedy installs a reading-and-accent channel into a released model, but it needs many recordings. This paper removes the recordings: the frozen backbone reads a sentence containing a common word it already says correctly, and that output serves as the teacher for the same sentence with the word replaced by an annotated reading with its pitch accent. Screened raters judged the tag right on 0.80 to 0.93 of unseen difficult words on four backbones spanning autoregressive, diffusion, and encoder-decoder synthesis; plain kana, which cannot express an accent, got 0.38 to 0.60. On words needing no edit, naturalness is non-inferior on one backbone; on the other three, listeners prefer the unedited rendition by 0.19 to 0.26.
发表机构
- KOWRO Inc.(KOWRO公司)
机构由 AI 辅助整理,请以论文原文为准。