arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10812cs.CLcs.AI

面向多语言机器翻译的开源大语言模型无参考后训练

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

Chris Han, Pengzhi Gao, Pei Fu, Jian Luan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多语言机器翻译,以开源大语言模型为对象,采用GRPO算法结合无参考质量评估奖励,通过检查点插值得到MiLMMT-46-v1.0,其翻译质量优于SFT模型及多款开源基线,在无参考评分上领先于谷歌翻译等专有系统,还发布了相关模型与代码。

中文摘要 AI 辅助

我们研究基于开源大语言模型的多语言机器翻译无参考后训练。从经监督微调的MiLMMT-46-v0.1模型出发,我们应用组相对策略优化(GRPO),其奖励函数为两个无参考质量评估模型的平均值,并通过语言识别进行门控。随后,我们对监督微调(SFT)和强化学习(RL)模型的检查点进行线性插值,得到MiLMMT-46-v1.0。在46种语言上,所得模型的翻译质量始终优于其SFT对应模型,表现优于近期的强大开源基线模型,包括Seed-X、HY-MT2和TranslateGemma,并在无参考评分上超越了评估的专有系统,如谷歌翻译、Gemini 3 Pro和GPT-5。我们进一步研究了同策略蒸馏,发现其达到了但未超越检查点插值RL所实现的质量前沿。我们发布了模型和代码以促进未来研究。

英文摘要

We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0. Across 46 languages, the resulting models consistently improve translation quality over their SFT counterparts, outperform strong recent open baselines, including Seed-X, HY-MT2, and TranslateGemma, and achieve leading reference-free scores against evaluated proprietary systems such as Google Translate, Gemini 3 Pro, and GPT-5. We further investigate on-policy distillation (OPD) and find that it reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation. We release the models and code to facilitate future research.

发表机构

  • Xiaomi Inc.(小米公司)

机构由 AI 辅助整理,请以论文原文为准。

↑