发表机构
Queen’s University Belfast; University of Technology Nuremberg; Saarland University; University of Melbourne; University of Helsinki; IIIT Hyderabad; TCS Research; Johns Hopkins University; Imam Ja’afar Al-Sadiq University; Cohere; Last Token; Charles University; University of Edinburgh; ETH Zurich; University of Zurich(贝尔法斯特女王大学; 纽伦堡工业大学; 萨尔大学; 墨尔本大学; 赫尔辛基大学; 海得拉巴国际信息技术学院; 塔塔咨询服务公司研究院; 约翰斯·霍普金斯大学; 伊玛目贾法尔·萨迪克大学; Cohere公司; Last Token公司; 查理大学; 爱丁堡大学; 苏黎世联邦理工学院; 苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多语言翻译基准的污染与区域忽视问题,提出基于FLORES的本地化基准Cultivar,测试32个开放权重模型,发现MT专用模型鲁棒性低、部分模型过拟合FLORES、模型翻译美国内容更优。
AI 中文摘要
多语言翻译基准通常以英语为源语言并翻译为其他语言,将语言对作为评估单元,这种设计易随时间产生污染,且忽视区域和文化因素。因此,我们提倡源对比评估,并通过Cultivar(FLORES的本地化子集,支持区域特定的翻译评估)实现该评估。当与未本地化的对应基准配对时,性能差异可用于探测数据污染和本地化鲁棒性。我们对32个开放权重模型进行基准测试,发现专门用于机器翻译(MT)的模型鲁棒性较低,部分模型可能过拟合FLORES,且无论语言如何,模型倾向于翻译美国内容的效果优于其他区域的内容。
英文摘要
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.