arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向ASR的基于音素的TTS数据增强:统一流水线与受控研究

Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation

Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang, Wei Xu

arXiv 2608.26697首次发表:更新:

发表机构

Shanghai Qi Zhi Institute; Megatronix (Beijing) Technology Co., Ltd.(上海期智研究院; 美格创芯(北京)科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于音素的TTS到ASR增强统一流水线及PFGS方法,经多语言ASR实验验证,PFGS可显著降低词错误率,明确了合成规模等关键控制变量。

AI 中文摘要

合成语音可为自动语音识别(ASR)提供可扩展的监督信号,但其效果取决于所选文本、参考语音及合成数据的规模。本文提出一种基于音素的统一TTS到ASR增强流水线,该流水线围绕采用F5-TTS架构、经从零开始训练且带语言ID条件的多语言TTS模型构建,融合了特定语言的字素到音素转换、参考语音过滤、候选文本选择、语音合成及匹配的ASR后续处理环节。本文进一步提出音素频率引导选择(PFGS)方法,该方法利用从真实ASR训练标签估算的音素频率对候选句子排序。针对阿拉伯语、法语、意大利语、葡萄牙语的独立单语ASR系统开展实验,覆盖13个测试集。在合成规模的遍历实验中,随机增强在11个测试集上的表现优于仅使用匹配真实数据的后续处理。在60%的标称合成预算下,PFGS在12个测试集上的表现优于仅使用真实数据的训练,在9个测试集上的表现优于随机选择;其相较随机选择的最大相对词错误率(WER)降低幅度达19.3%。在目标文本和合成数量固定的情况下,参考语音过滤分别使意大利语、法语Common Voice的绝对WER降低0.29和0.59个百分点。这些结果表明,合成规模、候选文本内容及参考质量是基于TTS的ASR增强中重要的控制变量。

英文摘要

Synthetic speech can provide additional supervision for automatic speech recognition (ASR), but constructing useful synthetic training data requires choosing both what to synthesize and how to synthesize it. We present a phoneme-guided text-to-speech (TTS) augmentation pipeline for ASR that connects multilingual speech generation with candidate-text selection and reference-speech quality control. Within this pipeline, we propose phoneme-frequency-guided selection (PFGS), which uses phoneme frequencies from real ASR training transcripts to prioritize candidate texts containing common phonetic content. Experiments with separate monolingual ASR systems cover four languages and 13 test sets. With random text selection, the pipeline improves recognition on 11 test sets at one or more synthesis ratios. PFGS further outperforms random selection on nine test sets, with relative word error rate (WER) reductions of up to 19.3%. An ablation with fixed target texts and synthesis counts further shows the benefit of reference-speech filtering. These results support using real-data phoneme statistics to guide the construction of effective synthetic supervision for ASR.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑