arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向发音的日语文本到语音强化学习:基于假名域ASR奖励

Pronunciation-Oriented Reinforcement Learning for Japanese Text-to-Speech with Kana-Domain ASR Rewards

Shiao Zhu, Lianbo Liu, Kai Washizaki, Koki Nikaido, Yui Sudo

arXiv 2610.07575首次发表:更新:

发表机构

SB Intuitions Corp.(SB Intuitions 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对日语TTS,提出基于假名域ASR奖励的Kana-CER强化学习方法,在GRPO下显著降低汉字读音错误并加速收敛,同时保持语音质量,且KL正则化可抑制输出延长。

AI 中文摘要

字符错误率(CER)由自动语音识别(ASR)计算,广泛用作文本到语音(TTS)系统强化学习(RL)后训练的可懂度奖励。然而,对于日语,正字法CER在面向发音的优化中引入了表示不匹配:不同的汉字读音可能映射到相同的正字法表示,而等效的发音可能采用不同的正字法形式。我们转而使用假名转写ASR模型和参考读音在假名域中计算CER(Kana-CER)。在匹配组相对策略优化(GRPO)条件下,与正字法CER奖励相比,Kana-CER将目标汉字读音错误相对减少约26%,同时保持相当的正字法CER和相似的说话人相似度及客观语音质量。它还以显著更少的优化步骤(4k对18k)达到其最佳验证性能。我们进一步观察到在无正则化的Kana-CER优化下出现严重的输出延长,而KL正则化可大幅抑制该现象。

英文摘要

Character error rate (CER) computed by automatic speech recognition (ASR) is widely used as an intelligibility reward for reinforcement learning (RL) post-training of text-to-speech (TTS) systems. For Japanese, however, orthographic CER introduces a representation mismatch for pronunciation-oriented optimization: distinct kanji readings may collapse to the same orthographic representation, while equivalent pronunciations may admit different orthographic forms. We instead compute CER in the kana domain using a kana-transcribing ASR model and reference readings (Kana-CER). Under matched group relative policy optimization (GRPO) conditions, Kana-CER reduces target-kanji reading error by approximately 26% relative to the orthographic CER reward, while maintaining comparable orthographic CER and similar speaker similarity and objective speech quality. It also reaches its best validation performance in substantially fewer optimization steps (4k vs. 18k). We further observe severe output elongation under unregularized Kana-CER optimization, which is substantially suppressed by KL regularization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑