无需训练的发音转写:基于文本约束的声学重打分
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
浏览论文内容
中文总结 AI 辅助
提出无需训练的ST2P流水线,融合词汇与声学信息,通过贪心搜索重打分,在日语等多语种上显著降低CER并提升速度。
中文摘要 AI 辅助
准确高效的发音转写对于大规模准备文本到语音训练数据至关重要。现有方法各有局限:字形到发音(G2P)和语音到发音(S2P)方法仅利用文本或仅利用语音,各自捕获部分信息;而语音和文本到发音(ST2P)方法虽同时使用两者,但需要昂贵的发音标注数据。为解决此问题,我们提出了一种无需训练的ST2P流水线,在推理时整合词汇和声学信息。词汇资源和G2P工具生成文本约束的候选,自左向右的贪心搜索利用冻结的预训练S2P模型的整序列负对数似然选择最佳候选。在三个日语语料库上,我们的方法将字符错误率(CER)从0.60%--1.40%(仅文本基线)降至0.04%--0.17%(使用参考转写),以及0.64%--1.58%(使用ASR转写)。它优于所有基线,包括一个训练的ST2P模型和商业多模态大语言模型。我们的贪心搜索方法在相似CER下比束搜索快3--3.5倍,级联比直接解码快2倍,确保了效率和准确性。在西班牙语、法语和初步英语中,它也超越了四个开放多模态大语言模型和最佳传统方法。
英文摘要
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.
发表机构
- Sakana AI
机构由 AI 辅助整理,请以论文原文为准。