arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27205cs.CLcs.AI

用户生成文本的音素化:基准、分类体系与组合方法

Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach

MinJu Jeon, Younghan Park, Han Sung Park, Jong-Hwan Kim, Dong-Jin Kim, Hoyeon Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对用户生成文本音素化难题,提出首个多语言基准UGTPhon及组合式G2P方法,通过显式规范形式建模显著降低非规范错误,小模型性能媲美大模型。

中文摘要 AI 辅助

语音合成系统越来越多地处理用户生成文本(UGT),例如“ppl”和“imo”,其发音必须从规范形式而非表面形式推断。我们引入了UGTPhon,这是首个针对英语、越南语和韩语UGT的字素到音素(G2P)基准,并附带一个基于推断的分类体系,用于细粒度诊断。现有的G2P模型和前沿大语言模型表现出系统性的规范到非规范性能差距,差距最高可达66.8个PER百分点。作为基准基线,我们提出了一种简单的组合式G2P方法,通过精确匹配查找和分阶段解码来整合规范形式证据。在匹配的ByT5和Qwen2.5-0.5B骨干网络上,显式规范形式建模持续减少非规范G2P错误。0.5B变体还与规模大得多的少样本前沿大语言模型表现相当,凸显了显式建模规范形式推断对UGT音素化的益处。

英文摘要

Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the first grapheme-to-phoneme (G2P) benchmark for UGT in English, Vietnamese, and Korean, together with an inference-grounded taxonomy for fine-grained diagnosis. Existing G2P models and frontier LLMs exhibit a systematic canonical-to-non-canonical performance gap, reaching up to 66.8 PER points. As a benchmark baseline, we propose a simple compositional G2P approach that incorporates canonical-form evidence through exact-match lookup and staged decoding. Across matched ByT5 and Qwen2.5-0.5B backbones, explicit canonical-form modeling consistently reduces non-canonical G2P errors. The 0.5B variant also performs competitively with much larger few-shot frontier LLMs, highlighting the benefit of explicitly modeling canonical-form inference for UGT phonemization.

发表机构

  • NAVER Cloud(NAVER云)
  • Hanyang University(汉阳大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑