发表机构
De La Salle University(德拉萨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对菲律宾语G2P中重音预测需句子级上下文的问题,提出用LLM辅助标注数据微调ByT5模型,显著降低PER和CER,并有效支持同形词消歧。
AI 中文摘要
字素到音素转换(G2P)是指将字素序列转换为相应音素序列的任务。尽管由于菲律宾语浅层正字法,其G2P相对简单,但重音等韵律特征的加入增加了复杂性,需要句子级上下文而非单词级输入。然而,菲律宾语的句子级数据通常不包含音素转写,这给训练G2P模型带来了挑战。因此,我们研究如何利用现有数据获取菲律宾语的句子级音素数据,并将所得模型与多语言单词级G2P进行比较,同时衡量它们预测菲律宾语重音标记位置的准确性。我们提出在基于ByT5的模型上进行微调,该模型预训练于多语言单词级G2P数据,并在三个由维基词典数据引导的LLM辅助流水线标注的句子级G2P数据集上进行微调。该方法产生的模型在G2P任务上表现良好,在人工校正的测试集上,最佳PER约为0.54%,CER约为2.50%,相比基础模型约19.74%的PER显著下降。该模型能够正确分类菲律宾语中大多数主要重音类别,但在处理malumi词时尤其困难。我们表明,基于ByT5的模型在句子级菲律宾语G2P上表现良好,并为菲律宾语同形词消歧提供了强大潜力。
英文摘要
Grapheme-to-phoneme conversion (G2P) refers to the task of converting a sequence of graphemes to a corresponding sequence of phonemes. While Filipino G2P is fairly straightforward due to its shallow orthography, the inclusion of prosodic features such as stress adds a layer of complexity that requires sentence-level context instead of single-word inputs. However, sentence-level data for Filipino typically do not include phoneme transcriptions, posing a challenge for training G2P models. As such, we investigate how to obtain sentence-level phoneme data for Filipino using available data and compare the resulting models with multilingual word-level G2P as well as measure how accurately they predict stress marker position for Filipino. We propose fine-tuning a ByT5-based model, pre-trained on multilingual word-level G2P data, on three sentence-level G2P datasets annotated with an LLM-assisted pipeline guided by data from Wiktionary. This approach produces models that perform well on the G2P task, achieving at best around 0.54% PER and 2.50% CER, a significant decrease compared to base model PER at around 19.74%, on a manually-corrected test set. The model is able to correctly classify most of the main stress classes in Filipino, but struggles particularly with malumi words. We show that a ByT5-based model performs well at sentence-level Filipino G2P and offers strong potential for Filipino homograph disambiguation.
CommentsAccepted at the 10th International Conference on Natural Language Processing and Information Retrieval (NLPIR 2026), Nara, Japan