推理的代价:神经机器翻译强化学习中的成本-质量权衡
The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation
浏览论文内容
中文总结 AI 辅助
研究神经机器翻译强化学习中推理痕迹对翻译质量的影响,通过在训练或推理阶段省略推理痕迹进行实验,发现推理尤其是推理阶段能提升质量,还研究了推理导致计算需求增加与翻译质量提升间的成本-质量权衡。
中文摘要 AI 辅助
具有可验证奖励的强化学习(RLVR)已成为大语言模型(LLMs)包括神经机器翻译(NMT)等下游任务训练后的可行范式。最新研究表明RLVR因诱导推理能力可能是翻译法律文件的首选训练方法,这引发其是否真归因于推理或更普遍的训练范式的问题。我们通过在训练或推理阶段系统地省略模型推理痕迹来研究其重要性。实验表明包含推理尤其是在推理阶段对整体翻译质量有积极影响。此外,推理导致输出令牌增加,因此我们研究了增加的计算需求与提高的翻译质量之间的成本-质量权衡。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the latest research indicating that RLVR could be the preferred training method for translating legal documents due to the induced reasoning capabilities, it raises the question whether it is really attributed to the reasoning or more generally to the training paradigm. We investigate the importance of including the model's reasoning trace in the generated responses during both training and inference by systematically omitting it from one of the phases. Our experiments show that including the reasoning, specifically during inference, has a positive effect on the overall translation quality. Furthermore, we recognise that the reasoning leads to an increase in output tokens, hence we study the cost-quality tradeoff between the increased computational demands and the improved translation quality.