阿拉伯语-俄语机器翻译基准测试:丰富形态与低词汇重叠条件下微调NMT与少样本LLM的比较
Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap
- Kazan Federal University(喀山联邦大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究针对阿拉伯语-俄语低资源翻译,对比微调NMT与少样本LLM,发现微调NLLB-1.3B以BLEU 16.3领先,显著优于LLM,低词汇重叠为主要挑战。
中文摘要 AI 辅助
阿拉伯语-俄语机器翻译(MT)由于阿拉伯语丰富的形态以及两种语言之间较低的词汇重叠度,至今仍研究不足。我们在一个新构建的1547万对语料库的20k/5k/5k划分上,对七个微调的神经机器翻译(NMT)模型与四个少样本大语言模型(LLM)进行了基准测试。微调的NLLB-1.3B取得了最高的BLEU(16.3)和COMET(0.738)分数。Aya-Expanse 8B在少样本LLM中领先(在500个句子上BLEU为1.7,chrF为25.7),但所有LLM的分数仍远低于微调的NMT基线。错误分析表明,低词汇重叠是主要的失败模式;在最差的翻译中,mT5-small产生了32%的过短输出。Bootstrap检验证实了大多数模型之间存在显著差异。我们的结果表明,在低资源条件下,微调的NMT在阿拉伯语-俄语翻译中显著优于少样本LLM。
英文摘要
Arabic-Russian machine translation (MT) remains under-explored due to the rich morphology of Arabic and low lexical overlap between the two languages. We benchmark seven fine-tuned neural machine translation (NMT) models against four few-shot large language models (LLMs) on a 20k/5k/5k split of a new 15.47M-pair corpus. Fine-tuned NLLB-1.3B achieves the highest BLEU (16.3) and COMET (0.738). Aya-Expanse 8B leads the few-shot LLMs (BLEU 1.7 on 500 sentences, chrF 25.7), but all LLM scores remain far below the fine-tuned NMT baselines. Error analysis identifies low lexical overlap as the dominant failure mode; among the worst translations, mT5-small produces 32% too-short outputs. Bootstrap tests confirm significant differences among most models. Our results demonstrate that fine-tuned NMT significantly outperforms few-shot LLMs for Arabic-Russian translation under low-resource conditions.