arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31727cs.CL

Ongiini-Eval-OW基准:关于Oshindonga和Oshikwanyama机器翻译与大语言模型计划基准的概念论文

The Ongiini-Eval-OW Benchmark: A Concept Paper for the Planned Benchmarking of Machine Translation and Large Language Models on Oshindonga and Oshikwanyama

Sebastian Küpers

首次发表
浏览论文内容

中文总结 AI 辅助

针对Oshiwambo语系中Oshindonga和Oshikwanyama方言缺乏机器翻译评估基准的问题,提出包含600条目的Ongiini-Eval-OW基准,涵盖多译者参考、现象分层和可复现评分,计划于2026年发布。

中文摘要 AI 辅助

Oshiwambo语——一个由超过一百万人使用的、彼此可理解的班图语族语言群,分布在纳米比亚北部和安哥拉南部,并且是大约半数纳米比亚家庭的家庭语言——据我们所知,尚无已发表的机器翻译评估基准。主要的商业服务(谷歌翻译、DeepL、微软翻译)、开放的多语言机器翻译模型(NLLB-200、MADLAD-400)以及开放的Masakhane检查点集合均未覆盖Oshindonga或Oshikwanyama这两种标准化方言中的任何一种。我们宣布Ongiini-Eval-OW,一个计划中的包含600个条目的英语-Oshindonga和英语-Oshikwanyama基准,其参考译文由两位独立译者以母语者身份提供,包含一个30条目的译者间一致性测试集,一个按11个标签进行现象标注的分层(每个标签至少30个条目),一个确定性的30%盲测划分,以及一个基于chrF++、BLEU和COMET-22的可复现评分协议,并辅以50条目的人工评估环节。我们记录了经验覆盖缺口、数据集构成、跨美国、欧洲和中国前沿及开放权重系统的启动排行榜模型矩阵,以及贡献流程。该数据集计划于2026年第四季度首次公开发布;本v1.0概念论文宣布了设计并呼吁参与。数据和代码将分别以CC-BY-4.0和MIT许可发布。

英文摘要

Oshiwambo -- a cluster of mutually intelligible Bantu languages spoken by over a million people across northern Namibia and southern Angola, and the home language of roughly half of Namibian households -- has, to our knowledge, no published machine-translation evaluation benchmark. Major commercial services (Google Translate, DeepL, Microsoft Translator), open multilingual MT models (NLLB-200, MADLAD-400), and the open Masakhane checkpoint collection all lack coverage of either standardised dialect, Oshindonga or Oshikwanyama. We announce Ongiini-Eval-OW, a planned 600-item English-Oshindonga and English-Oshikwanyama benchmark with native-speaker references from two independent translators, a 30-item inter-translator agreement set, an 11-tag phenomenon-tagged stratification (at least 30 items per tag), a deterministic 30% blind split, and a reproducible scoring protocol over chrF++, BLEU, and COMET-22, supplemented by a 50-item human-evaluation round. We document the empirical coverage gap, the dataset composition, the launch-leaderboard model matrix across American, European, and Chinese frontier and open-weight systems, and the contribution pipeline. The dataset is targeted for first public release in Q4 2026; this v1.0 concept paper announces the design and the call for participation. Data and code will be released under CC-BY-4.0 and MIT respectively.

补充信息

↑