DRT:基于长思维链的深度推理翻译
DRT: Deep Reasoning Translation via Long Chain-of-Thought
- Pattern Recognition Center, WeChat AI, Tencent Inc(模式识别中心、微信AI、腾讯公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对文学文本中明喻隐喻因文化差异难翻译的问题,提出DRT模型,通过挖掘相关句子构建多智能体长思维翻译框架并生成训练数据,以Qwen2.5和LLama-3.1为骨干,性能优于普通及仅配对句微调的LLMs。
AI中文摘要:
近年来,O1类模型已成为代表性案例,展示了长思维链(CoT)在数学和编码等推理任务中的有效性。本文提出DRT,尝试将长CoT的成功应用于神经机器翻译(MT)。具体而言,考虑到可能包含明喻和隐喻的文学书籍,由于文化差异,将这些文本翻译成目标语言在实践中非常困难。在这种情况下,直译往往无法有效传达预期含义。即使对于专业的人工译者,在整个翻译过程中也必须仔细考虑如何保留语义。为了在MT中模拟大语言模型(LLMs)的长思维能力,我们首先从现有文学书籍中挖掘包含明喻或隐喻的句子,然后开发一个多智能体框架,通过长思维来翻译这些句子。在该多智能体框架中,翻译器在顾问提供的建议下迭代翻译源句。为了确保长思维的有效性,还使用评估器量化每一轮的翻译质量。通过这种方式,我们收集了数万条长思维MT数据,用于训练我们的DRT。以Qwen2.5和LLama-3.1为骨干,DRT模型能够学习机器翻译过程中的思维过程,其性能优于普通LLMs以及仅在无长思维的配对句子上微调的LLMs,证明了其有效性。合成数据和模型检查点已在https://github.com/krystalan/DRT发布。
英文摘要:
Recently, O1-like models have emerged as representative examples, illustrating the effectiveness of long chain-of-thought (CoT) in reasoning tasks such as math and coding tasks. In this paper, we introduce DRT, an attempt to bring the success of long CoT to neural machine translation (MT). Specifically, in view of the literature books that might involve similes and metaphors, translating these texts to a target language is very difficult in practice due to cultural differences. In such cases, literal translation often fails to convey the intended meaning effectively. Even for professional human translators, considerable thought must be given to preserving semantics throughout the translation process. To simulate LLMs' long thought ability in MT, we first mine sentences containing similes or metaphors from existing literature books, and then develop a multi-agent framework to translate these sentences via long thought. In the multi-agent framework, a translator is used to iteratively translate the source sentence under the suggestions provided by an advisor. To ensure the effectiveness of the long thoughts, an evaluator is also employed to quantify the translation quality in each round. In this way, we collect tens of thousands of long-thought MT data, which is used to train our DRT. Using Qwen2.5 and LLama-3.1 as the backbones, DRT models can learn the thought process during machine translation, and outperform vanilla LLMs as well as LLMs which are simply fine-tuning on the paired sentences without long thought, showing its effectiveness. The synthesized data and model checkpoints are released at https://github.com/krystalan/DRT.