EnSiTa——面向领域特定机器翻译的三语多领域平行数据集与基准
EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation
浏览论文内容
中文总结 AI 辅助
针对低资源语言专业领域平行数据稀缺问题,构建三语多领域数据集EnSiTa,系统评估多种模型与设置,为领域特定MT提供最全面的基准。
中文摘要 AI 辅助
低资源语言的机器翻译(MT)仍远落后于高资源语言,且在专业领域差距最大,因为平行数据稀缺或完全缺失。我们提出EnSiTa,一个面向英语、僧伽罗语和泰米尔语的三语多领域平行数据集与基准。EnSiTa提供由专业译员在多年严格质量控制流程下制作的人工后期编辑训练数据,覆盖七个领域,以及针对这些领域及一个额外领域的测试集。利用该数据集,我们对所有六个语言方向进行了领域特定MT的广泛研究,微调了从零训练的Transformer、预训练翻译模型(NLLB-600M)以及仅解码器LLM(Gemma 3系列,1B-12B,及TranslateGemma),涵盖不同训练数据规模、模型规模以及域内、跨域、多语言和多领域设置。据我们所知,这是低资源MT领域最广泛的系统化多领域平行数据构建与基准测试工作。我们的数据和模型将公开发布。
英文摘要
Machine Translation (MT) for low-resource languages remains far behind that of high-resource languages, and the gap is widest in specialised domains, where parallel data is scarce or entirely absent. We present EnSiTa, a trilingual multi-domain parallel dataset and benchmark for English, Sinhala and Tamil. EnSiTa provides human post-edited training data for seven domains, plus manually translated test sets for those and one additional domain, all produced by professional translators under a multi-year, rigorously quality-controlled process. Using this dataset, we conduct an extensive study of domain-specific MT for all six language directions, fine-tuning a from-scratch Transformer, a pre-trained translation model (NLLB-600M), and decoder-only LLMs (Gemma 3 family, 1B-12B, and TranslateGemma) across training-data sizes, model scales, and in-domain, cross-domain, multilingual and multi-domain settings. To the best of our knowledge, this is the most extensive systematically documented multi-domain parallel data creation and benchmarking effort for low-resource MT. Our data and models will be publicly released.
发表机构
- Massey University(梅西大学)
- University of Moratuwa(莫拉图瓦大学)
- National University of Singapore(新加坡国立大学)
- Utrecht University(乌得勒支大学)
- Rowan University(罗文大学)
- District General Hospital, Hambantota(汉班托塔地区综合医院)
- WSO2
机构由 AI 辅助整理,请以论文原文为准。