arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用语法书进行低资源机器翻译合成数据生成的析因研究

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

Varun Ghat Ravikumar, Sina Ahmadi, Lena Jäger, Rico Sennrich

arXiv 2607.22376首次发表:更新:

发表机构

University of Zurich(苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对濒危语言机器翻译缺平行数据问题,利用大语言模型从语法书提取内容生成合成语料库微调,经三种低资源语言验证,通过析因研究确定增益因素组合,证明可将静态语言文档用于机器翻译微调,为资源匮乏语言提供翻译工具路径。

AI 中文摘要

尽管存在描述性语法书,但大多数濒危语言缺乏机器翻译所需的平行数据。我们引入了一种管道,利用大语言模型从语法书中提取语法规则、例句和词汇表,并生成合成平行语料库用于微调,而不是像先前工作那样在推理时将语法内容输入提示中。在三种类型不同的低资源语言——卡拉芒语(巴布亚语系)、图阿钦语(罗曼语族)和曼丹语(苏语族)上进行验证,结果表明,在75%的卡拉芒语配置和59%的图阿钦语配置中,基于合成数据的微调优于种子数据基线,最佳情况下ChrF++增益分别为+8.8、+5.3和+3.3。通过对96种配置进行系统的析因研究,我们确定了哪些因素组合能带来增益以及它们在何处失效。我们的结果表明,静态语言文档可重新用于机器翻译微调,为资源严重不足的语言提供了实用的翻译工具路径。

英文摘要

Most endangered languages lack the parallel data required for machine translation, despite the existence of descriptive grammar books. We introduce a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work. Validated on three typologically diverse low-resource languages-Kalamang (Papuan), Tuatschin (Romance), and Mandan (Siouan)-we show that fine-tuning on synthetic data improves over seed-data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, with best-case ChrF++ gains of +8.8, +5.3, and +3.3 respectively. Through a systematic factorial study across 96 configurations varying target part-of-speech, retrieval granularity, and sample volume, we identify which factor combinations drive gains and where they break down. Our results demonstrate that static linguistic documentation can be repurposed for machine translation fine-tuning, offering a practical path towards translation tools for severely under-resourced languages.

CommentsAccepted at CLiC-it 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑