arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.11919cs.CL

LLM推理用于机器翻译:基于思考令牌的合成数据生成

LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens

  • Inria Paris(巴黎研究所)

机构由 AI 辅助整理,请以论文原文为准。

Armel Zebaze, Rachel Bawden, Benoît Sagot

更新

AI总结:

本研究探讨大型推理模型在机器翻译中使用思考令牌的效果,发现其无益,但结合翻译特定提示的中间令牌可提升性能,建议教师模型改进翻译或扩展语料库更有效。

AI中文摘要:

大型推理模型(LRMs)通过在回答查询之前构思自然语言思维过程,为问题解决带来了新的可能性。尽管它们在数学和编程任务中的能力众所周知,但它们对机器翻译(MT)任务的影响仍未得到充分探索。在本工作中,我们探讨了在多种资源丰富程度不同的语言对和多种设置下执行MT时生成中间令牌的益处。我们发现“思考令牌”并不能帮助LRMs更好地执行MT。这一结果也推广到使用蒸馏思维链(CoT)进行微调的模型,这些CoT受人类翻译者实践的启发,在翻译前进行推理。具体而言,使用合成的CoT解释(详细说明如何逐步翻译)对模型进行微调,并不优于标准的输入-输出微调。然而,通过结合模块化翻译特定提示策略的输出构建中间令牌,则带来了改进。我们的发现强调了微调过程中中间令牌的贡献高度依赖于其中是否包含翻译尝试。更广泛地说,我们的结果表明,使用教师模型来改进目标翻译或扩展平行语料库,比将它们的CoT解释蒸馏到“思考”MT模型中更有影响力。

英文摘要:

Large reasoning models (LRMs) have led to new possibilities in terms of problem-solving, through the devising of a natural language thought process prior to answering a query. While their capabilities are well known across mathematics and coding tasks, their impact on the task of machine translation (MT) remains underexplored. In this work, we explore the benefits of the generation of intermediate tokens when performing MT across multiple language pairs of different levels of resourcedness and multiple setups. We find that "thinking tokens" do not help LRMs better perform MT. This result generalizes to models fine-tuned to reason before translating using distilled chain of thought (CoT) inspired by human translators' practices. Specifically, fine-tuning a model with synthetic CoT explanations detailing how to translate step-by-step does not outperform standard input-output fine-tuning. However, constructing the intermediate tokens by combining the outputs of modular translation-specific prompting strategies results in improvements. Our findings underscore that the contribution of intermediate tokens during fine-tuning highly depends on the presence of translation attempts within them. More broadly, our results suggest that using a teacher to refine target translations or to expand parallel corpora is more impactful than distilling their CoT explanations into "thinking" MT models.

↑