arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用统计机器学习实现英语与叙利亚语(东部叙利亚方言)之间的机器翻译

Machine Translation between English and Syriac (East Syriac Dialect) using Statistical Machine Learning

Hadiana Sliwa, Hossein Hassani

arXiv 2609.18529首次发表:更新:

发表机构

University of Kurdistan Hewlêr(库尔德斯坦赫勒尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对濒危的叙利亚语,利用Moses框架构建首个基于短语的英语-叙利亚语统计机器翻译模型,创建38,847句对数据集,最佳BLEU为23.54,并公开资源以推动该语言NLP研究。

AI 中文摘要

联合国教科文组织将亚述语(叙利亚语)视为濒危语言。尽管亚述人在世界各地使用该语言,但使用人口数量不确定(范围在50万至150万之间)。叙利亚语也是自然语言处理(NLP)中研究最少的语言之一。尽管过去十年机器翻译(MT)取得了进步,但缺乏公开可用的语料库以及叙利亚文字(特别是Madnkhaya文字)的正字法复杂性,使得该语言在计算语言学文献中完全被忽视。本研究使用Moses框架开发了第一个基于短语的统计机器翻译(SMT)模型,用于英语到亚述语的翻译。我们从完整的英语和叙利亚语圣经中创建了一个包含38,847个句对的数据集,将已有的新约数据集与通过PDF提取从头构建的旧约数据集合并,使用了自定义分割脚本和三位双语标注者的人工对齐审查。语料库的叙利亚语部分在训练前经过去除变音符号和字节对编码分词处理,以减少正字法稀疏性。我们使用不同的配置和分割方案比例、语言模型阶数、扭曲限制以及是否包含操作序列模型,训练并评估了六个模型。最佳配置的词级BLEU得分为23.54。由11位亚述语母语者进行的人工评估得出,平均充分性和流利度得分分别为3.42和3.34(满分5分)。这些结果与针对形态丰富的闪米特语言在圣经语料库上训练的可比低资源SMT模型一致。语料库、脚本和训练好的模型均已公开,为研究社区提供了第一个系统整理的英语-叙利亚语数据集,并为未来针对这种濒危语言的机器翻译和更广泛的NLP工作提供了可复现的基线。

英文摘要

UNESCO considers the Assyrian (Syriac) language an endangered language. Although Assyrians speak the language worldwide, the speaking population is uncertain (ranging from 500,000 to 1,500,000). Syriac is also one of the least studied languages in Natural Language Processing (NLP). Despite advances in Machine Translation (MT) over the past decade, the lack of publicly available corpora and the orthographic complexity of the Syriac script, specifically the Madnkhaya script, have left this language entirely ignored in the computational linguistics literature. This study develops the first phrase-based Statistical MT (SMT) model for English-to-Assyrian MT using the Moses framework. We created a dataset of 38,847 sentence pairs from the complete English and Syriac Bible, merging a pre-existing New Testament dataset with an Old Testament built from scratch through PDF extraction, using custom segmentation scripts and manual alignment review by three bilingual annotators. The Syriac side of the corpus undergoes diacritic removal and Byte-Pair Encoding tokenization to reduce orthographic sparsity before training. We trained and evaluated six models using different configurations and splitting-scheme ratios, language model order, distortion limits, and the inclusion of an Operation Sequence Model. The best-performing configuration achieves a word-level BLEU score of 23.54. Human evaluation by 11 native Assyrian speakers resulted in mean adequacy and fluency scores of 3.42 and 3.34 out of 5, respectively. These results are consistent with comparable low-resource SMT models trained on Biblical corpora for morphologically rich Semitic languages. The corpora, scripts, and trained model are publicly available, providing the research community with the first systematically curated English-Syriac dataset and a reproducible baseline for future MT and broader NLP work on this endangered language.

Comments17 pages, 4 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑