AI 中文总结
研究法律领域神经机器翻译难题,通过比较多种方法,评估小型语言模型在不同再训练范式下的表现,以瑞士法律系统为测试平台,发现强化学习效果好,增强小型模型接近前沿推理模型,再训练范式随模型规模收益递减。
AI 中文摘要
法律领域的神经机器翻译是一项对语言和概念要求很高的任务,因其语言复杂性和高精度要求。具备推理能力的语言模型为应对此类挑战带来新可能。本文比较多种增强翻译质量的方法,评估如Qwen3.5 4B等小型语言模型在不同再训练范式下的效果,并与前沿推理模型对比。以瑞士法律系统为测试平台,结果表明小型“基础”模型质量可大幅提高,可验证奖励的强化学习在法律领域机器翻译中效果优于监督微调,增强的小型模型性能接近前沿推理模型但仍有差距,且再训练范式随模型规模增大收益递减。代码和模型可公开获取。
英文摘要
Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity of legal language and the high level of precision it requires. The recent emergence of reasoning-capable language models opens new possibilities for tackling such challenges. They add to a set of other previously proposed techniques to enhance the translation quality, which includes supervised fine-tuning and reinforcement learning. In this work, we perform a comparison between these various approaches. More particularly, we evaluate small language models such as Qwen3.5 4B, Qwen3.5 9B, and Gemma 3 12B enhanced with various re-training paradigms and compare their performances against frontier reasoning models. We focus on the Swiss legal system, which -- with its unique multilingual statutes -- offers a particularly challenging testbed for reasoning-augmented models. Our results show that the quality of small ``base'' models can be greatly enhanced, and that reinforcement learning with verifiable rewards can be applied to NMT in the legal domain and surpasses the translation quality of supervised fine-tuning. The performance of enhanced small models is close to the one of state-of-the-art reasoning models yet remains inferior. We also note that re-training paradigms yield diminishing returns as model size increase. The code and models are publicly available at https://github.com/aixiuxiuxiu/Legal-MT-SFT-RL.
CommentsCode available at https://github.com/aixiuxiuxiu/Legal-MT-SFT-RL