发表机构
National Institute of Technology Meghalaya(梅加拉亚国家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对资源匮乏的普纳尔语,构建首个英语-普纳尔语统计机器翻译系统,基于10234句平行语料训练,建立定量基准,分析错误并探讨未来方向。
AI 中文摘要
普纳尔语(Pnar)是一种使用人口约40万的南亚语系语言,分布于印度梅加拉亚邦贾因蒂亚丘陵地区,目前缺乏数字语料库和自然语言处理(NLP)资源。本文针对英语-普纳尔语这一语言对开展了首个机器翻译研究。利用从《Wyrta》报纸收集的文章,我们构建了包含10234个句子的平行语料库,并基于9563个平行语料,使用Moses、GIZA++、KenLM工具,针对每个翻译方向设置三种配置(采用不同的词汇化重排序策略和最小错误率训练(MERT)调优),训练了基于短语的统计机器翻译(SMT)系统。模型在包含371个句子的保留测试集上进行评估,表现最佳的系统在普纳尔语译英语任务中BLEU分数达14.97(chrF2:33.42,TER:77.60),在英语译普纳尔语任务中BLEU分数达11.16(chrF2:31.38,TER:93.51),为该语言对建立了首个定量基准。词汇化重排序使普纳尔语译英语的BLEU分数提升了3.73个百分点,这反映了源语言的SOV词序向目标语言SVO词序的结构转变;而在低资源条件下,MERT调优会降低BLEU性能。最后,我们分析了剩余的翻译错误,包括形态层面的未登录词(OOV)、长距离重排序以及卡西语(Khasi)代码混合,并探讨了针对普纳尔语的神经机器翻译和多语言机器翻译的未来研究方向。
英文摘要
Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural language processing (NLP) resources. This paper presents the first machine translation study for the English and Pnar language pair. Using articles collected from the Wyrta newspaper, we built a parallel corpus comprising of 10,234 sentences and trained phrase-based statistical machine translation (SMT) systems the models using 9,563 parallel corpora under three configurations for each direction using Moses, GIZA++ , KenLM, varying lexicalized reordering and minimum error rate training (MERT) tuning. The models are evaluated on a held out test set of 371 sentences, the best performing system achieves a BLEU score of 14.97 (chrF2: 33.42, TER: 77.60) for Pnar to English and 11.16 (chrF2: 31.38, TER: 93.51) for English to Pnar, establishing the first quantitative benchmark for this language pair. Lexicalized reordering improves translation quality by 3.73 BLEU points for Pnar to English, reflecting the structural shift from the source language's SOV word order to the target language's SVO order, whereas MERT tuning degrades BLEU performance under low resource conditions. Finally, we analyze the remaining translation errors, including morphological out of vocabulary (OOV) words, long-distance reordering and Khasi code mixing and discuss future directions toward neural and multilingual machine translation for Pnar.