arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

墨尔本大学WMT 2026 CreoleMT提交:面向低资源太平洋克里奥尔语机器翻译的领域平衡方法

The University of Melbourne WMT 2026 CreoleMT Submission: A Domain-Balanced Approach to Low-Resource Pacific Creole Machine Translation

Raphaël Merx, Nick Thieberger, Ekaterina Vylomova

arXiv 2609.13615首次发表:更新:

发表机构

The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对WMT26克里奥尔语翻译任务,采用预训练加领域平衡微调的方法,结合LLM辅助重写、反向翻译和蒸馏技术,在Bouquet及口语测试集上超越基线3+ chrF++点。

AI 中文摘要

针对我们提交至WMT26克里奥尔语翻译共享任务的工作,我们聚焦于太平洋克里奥尔语的机器翻译(MT)模型:托克皮辛语、比斯拉马语和所罗门皮钦语,并特别关注广泛领域的性能。在大量领域不平衡数据上进行预训练后,我们继续在多样化的领域平衡数据混合上进行微调。我们依赖多种数据收集和准备技术,包括大语言模型辅助的拼写重写和对齐、反向翻译,以及从Gemini蒸馏训练数据中原本不存在的领域。在Bouquet和一个由口语转录组成的新测试集上评估,我们的模型在所有方向上以人工原始参考超越开放模型基线3个以上的chrF++点。展望未来,我们计划为所罗门皮钦语和比斯拉马语开发人工翻译的测试集,并将我们最好的模型蒸馏成更小的模型,同时保留广泛的领域覆盖。

英文摘要

For our submission to the WMT26 Creole Language Translation Shared Task, we focus on machine translation (MT) models for Pacific creoles: Tok Pisin, Bislama, and Solomon Pijin, with particular attention to broad domain performance. After pre-training on a large collection of domain-imbalanced data, we continue fine-tuning on a diverse mix of domain-balanced data. We rely on a number of data collection and preparation techniques, including LLM-assisted respelling and alignment, back-translation, and distillation from Gemini for domains originally not present in training data. Evaluated on Bouquet and a novel test set made of spoken language transcripts, our models beat open model baselines by 3+ chrF++ points in all directions with human-original references. Looking ahead, we plan to develop human-translated test sets for Solomon Pijin and Bislama, and to distil our best models into much smaller ones that retain broad domain coverage.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑