arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19805cs.CL

字典约束的未分词语言字素到音素转换:基于LLM标注数据

Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data

  • Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

Rui Hu, Zhenpeng Zhan, Xiaolong Lin

AI总结:

本文提出一种基于字典约束CRF的上下文感知神经G2P方法,利用LLM生成超200万句子解决数据稀缺,在Joyo-Kanji-Yomi基准上显著优于传统方法。

AI中文摘要:

字素到音素(G2P)转换将原始文本转换为其音素形式,是文本到语音(TTS)和自动语音识别(ASR)系统的重要组成部分。它需要快速、稳定且具有上下文感知能力。对于日语等未分词语言,G2P还需将分词与高度依赖上下文的同音字消歧相结合,而准确标注数据的稀缺仍是瓶颈。本文提出了一种上下文感知的神经G2P方法,该方法对由字典构建的词格上的判别式条件随机场(CRF)路径进行评分。为解决数据稀缺问题,我们利用大语言模型(LLMs)生成了超过200万条句子。实验结果表明,我们的方法显著优于传统的基于形态分析器的方法和神经序列模型。在Joyo-Kanji-Yomi基准上,我们的方法达到了99.62%的目标词阅读准确率、0.32%的目标词音素错误率(PER)和0.14%的句子PER。

英文摘要:

Grapheme-to-phoneme (G2P) conversion turns raw text into its phonemic form and is an essential part of both text-to-speech (TTS) and automatic speech recognition (ASR) systems. It is required to be fast, stable and context-aware. For unsegmented languages such as Japanese, G2P additionally couples word segmentation with highly context-dependent polyphone disambiguation, and the scarcity of accurately annotated data remains a bottleneck. In this paper, we present a context-aware, segmentation-agnostic neural G2P framework that models the joint segmentation-and-reading hypothesis space, scoring paths of a discriminative conditional random field (CRF) over a dictionary-derived word lattice. To tackle data scarcity, we utilize large language models (LLMs) to generate more than 2 million sentences. Experimental results demonstrate that our method substantially outperforms conventional morphological analyzer-based methods and neural sequence models. On the Joyo-Kanji-Yomi benchmark, our method reaches 99.62% target word reading accuracy, 0.32% target word phoneme error rate (PER) and 0.14% sentence PER.

补充信息

↑