arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cultivar:用于研究污染与本地化鲁棒性的对比及面向区域的翻译基准

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao

arXiv 2608.09766首次发表:更新:

发表机构

Queen’s University Belfast; University of Technology Nuremberg; Saarland University; University of Melbourne; University of Helsinki; IIIT Hyderabad; TCS Research; Johns Hopkins University; Imam Ja’afar Al-Sadiq University; Cohere; Last Token; Charles University; University of Edinburgh; ETH Zurich; University of Zurich(贝尔法斯特女王大学; 纽伦堡工业大学; 萨尔大学; 墨尔本大学; 赫尔辛基大学; 海得拉巴国际信息技术学院; 塔塔咨询服务公司研究院; 约翰斯·霍普金斯大学; 伊玛目贾法尔·萨迪克大学; Cohere公司; Last Token公司; 查理大学; 爱丁堡大学; 苏黎世联邦理工学院; 苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多语言翻译基准的污染与区域忽视问题,提出基于FLORES的本地化基准Cultivar,测试32个开放权重模型,发现MT专用模型鲁棒性低、部分模型过拟合FLORES、模型翻译美国内容更优。

AI 中文摘要

多语言翻译基准通常以英语为源语言并翻译为其他语言,将语言对作为评估单元,这种设计易随时间产生污染,且忽视区域和文化因素。因此,我们提倡源对比评估,并通过Cultivar(FLORES的本地化子集,支持区域特定的翻译评估)实现该评估。当与未本地化的对应基准配对时,性能差异可用于探测数据污染和本地化鲁棒性。我们对32个开放权重模型进行基准测试,发现专门用于机器翻译(MT)的模型鲁棒性较低,部分模型可能过拟合FLORES,且无论语言如何,模型倾向于翻译美国内容的效果优于其他区域的内容。

英文摘要

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑