arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向重音感知的句子级菲律宾语G2P:基于弱监督ByT5微调

Towards Stress-Aware Sentence-Level Filipino G2P With Weakly-Supervised ByT5 Fine-Tuning

Lorenz Bernard Marqueses, Paulo Grane Gabriel Silva, Chastine Cabatay, Ericson Adler Tan, Ann Franchesca Laguna

arXiv 2609.09974首次发表:更新:

发表机构

De La Salle University(德拉萨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对菲律宾语G2P中重音预测需句子级上下文的问题,提出用LLM辅助标注数据微调ByT5模型,显著降低PER和CER,并有效支持同形词消歧。

AI 中文摘要

字素到音素转换(G2P)是指将字素序列转换为相应音素序列的任务。尽管由于菲律宾语浅层正字法,其G2P相对简单,但重音等韵律特征的加入增加了复杂性,需要句子级上下文而非单词级输入。然而,菲律宾语的句子级数据通常不包含音素转写,这给训练G2P模型带来了挑战。因此,我们研究如何利用现有数据获取菲律宾语的句子级音素数据,并将所得模型与多语言单词级G2P进行比较,同时衡量它们预测菲律宾语重音标记位置的准确性。我们提出在基于ByT5的模型上进行微调,该模型预训练于多语言单词级G2P数据,并在三个由维基词典数据引导的LLM辅助流水线标注的句子级G2P数据集上进行微调。该方法产生的模型在G2P任务上表现良好,在人工校正的测试集上,最佳PER约为0.54%,CER约为2.50%,相比基础模型约19.74%的PER显著下降。该模型能够正确分类菲律宾语中大多数主要重音类别,但在处理malumi词时尤其困难。我们表明,基于ByT5的模型在句子级菲律宾语G2P上表现良好,并为菲律宾语同形词消歧提供了强大潜力。

英文摘要

Grapheme-to-phoneme conversion (G2P) refers to the task of converting a sequence of graphemes to a corresponding sequence of phonemes. While Filipino G2P is fairly straightforward due to its shallow orthography, the inclusion of prosodic features such as stress adds a layer of complexity that requires sentence-level context instead of single-word inputs. However, sentence-level data for Filipino typically do not include phoneme transcriptions, posing a challenge for training G2P models. As such, we investigate how to obtain sentence-level phoneme data for Filipino using available data and compare the resulting models with multilingual word-level G2P as well as measure how accurately they predict stress marker position for Filipino. We propose fine-tuning a ByT5-based model, pre-trained on multilingual word-level G2P data, on three sentence-level G2P datasets annotated with an LLM-assisted pipeline guided by data from Wiktionary. This approach produces models that perform well on the G2P task, achieving at best around 0.54% PER and 2.50% CER, a significant decrease compared to base model PER at around 19.74%, on a manually-corrected test set. The model is able to correctly classify most of the main stress classes in Filipino, but struggles particularly with malumi words. We show that a ByT5-based model performs well at sentence-level Filipino G2P and offers strong potential for Filipino homograph disambiguation.

CommentsAccepted at the 10th International Conference on Natural Language Processing and Information Retrieval (NLPIR 2026), Nara, Japan

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑