发表机构
Faculdade de Engenharia da Universidade Lúrio; Universidade Eduardo Mondlane - Escola de Comunicação e Artes(卢里奥大学工程学院; 爱德华多·蒙德拉纳大学传播与艺术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究扩展FLORES+,添加三种莫桑比克班图语变体,发现参考译文变体混淆显著影响评估,变体感知微调能提升目标变体性能,并强调变体感知标识与报告的必要性。
AI 中文摘要
在本文中,我们通过为三种莫桑比克班图语变体(希昌加纳语、莫桑比克尼扬贾语和塞纳语)添加葡萄牙语源评估集,扩展了FLORES+。我们将希昌加纳语与现有的聪加语参考译文进行比较,将莫桑比克尼扬贾语与奇契瓦语进行比较,并评估了NLLB-200、谷歌翻译、GPT以及一个变体感知的NLLB模型。固定系统输出不变,揭示了显著的参考译文敏感性。在devtest上,仅将参考译文从聪加语改为希昌加纳语,就使NLLB-200的spBLEU降低了13.10分,谷歌翻译降低了15.30分。在匹配的尼扬贾语子集上,用莫桑比克尼扬贾语替换奇契瓦语,产生了较小但一致的降低,分别降低了3.03和6.10个spBLEU点。变体感知的微调在预期目标上逆转了这一模式:相对于NLLB-200,它在devtest上将希昌加纳语提高了7.04个spBLEU点,莫桑比克尼扬贾语提高了5.33个点,同时在兄弟参考译文上性能有所下降。GPT在聪加语和奇契瓦语上具有竞争力,但在莫桑比克变体上明显较弱。对于塞纳语,微调模型在devtest上达到了12.64的spBLEU和36.21的chrF++。这些发现促使针对跨境语言或语言方言/变体,采用变体感知的语言标识符、参考译文和报告。数据已在Hugging Face上公开,网址为https URL。
英文摘要
In this paper, we extend FLORES+ with Portuguese-source evaluation sets for three Mozambican Bantu varieties: Xichangana, Mozambican Nyanja, and Sena. We compare Xichangana with the existing Tsonga reference and Mozambican Nyanja with Chichewa, and evaluate NLLB-200, Google Translate, GPT, and a variant-aware NLLB model. Holding system output fixed reveals substantial reference sensitivity. On \textit{devtest}, changing only the reference from Tsonga to Xichangana reduces spBLEU by 13.10 points for NLLB-200 and 15.30 for Google. On matched Nyanja subsets, replacing Chichewa with Mozambican Nyanja produces smaller but consistent reductions of 3.03 and 6.10 spBLEU, respectively. Variant-aware fine-tuning reverses this pattern on the intended targets: relative to NLLB-200, it improves Xichangana by 7.04 spBLEU and Mozambican Nyanja by 5.33 on \textit{devtest}, while losing performance on the sibling references. GPT is competitive on Tsonga and Chichewa but substantially weaker on the Mozambican varieties. For Sena, the finetuned model reaches 12.64 spBLEU and 36.21 chrF++ on \textit{devtest}. These findings motivate variety-aware language identifiers, references, and reporting for cross-border languages or language dialects/variants. The data is publicly available on Hugging Face at https://huggingface.co/datasets/MOZNLP/FLORES_MOZ
CommentsAccepted at WMT2026