arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30924cs.CLcs.SDeess.AS

无需训练的发音转写:基于文本约束的声学重打分

Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring

Hikaru Asano, Yotaro Kubo, So Kuroki

首次发表
浏览论文内容

中文总结 AI 辅助

提出无需训练的ST2P流水线,融合词汇与声学信息,通过贪心搜索重打分,在日语等多语种上显著降低CER并提升速度。

中文摘要 AI 辅助

准确高效的发音转写对于大规模准备文本到语音训练数据至关重要。现有方法各有局限:字形到发音(G2P)和语音到发音(S2P)方法仅利用文本或仅利用语音,各自捕获部分信息;而语音和文本到发音(ST2P)方法虽同时使用两者,但需要昂贵的发音标注数据。为解决此问题,我们提出了一种无需训练的ST2P流水线,在推理时整合词汇和声学信息。词汇资源和G2P工具生成文本约束的候选,自左向右的贪心搜索利用冻结的预训练S2P模型的整序列负对数似然选择最佳候选。在三个日语语料库上,我们的方法将字符错误率(CER)从0.60%--1.40%(仅文本基线)降至0.04%--0.17%(使用参考转写),以及0.64%--1.58%(使用ASR转写)。它优于所有基线,包括一个训练的ST2P模型和商业多模态大语言模型。我们的贪心搜索方法在相似CER下比束搜索快3--3.5倍,级联比直接解码快2倍,确保了效率和准确性。在西班牙语、法语和初步英语中,它也超越了四个开放多模态大语言模型和最佳传统方法。

英文摘要

Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.

发表机构

  • Sakana AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑