arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

带思考的翻译:面向多领域机器翻译的难度自适应推理强化学习方法

Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation

Yongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen, Xiaodong Shi

arXiv 2607.29287首次发表:更新:

发表机构

Institute of Artificial Intelligence, Xiamen University; School of Informatics, Xiamen University; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism(厦门大学人工智能研究院; 厦门大学信息学院; 文化和旅游部闽台非物质文化遗产数字化保护与智能处理重点实验室(厦门大学))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多领域机器翻译的难度差异挑战,提出TwT框架,经两阶段训练后,其7B和14B参数版本在翻译质量上优于更大的SOTA模型,且token使用量降低32%-60%。

AI 中文摘要

多领域机器翻译(MDMT)因各领域语言复杂程度不同而面临独特挑战。受人类翻译员能根据难度调整推理精力的启发,我们提出TwT(Translation with Thought,带思考的翻译),这是一种资源理性框架,可学习在直觉推理与审慎推理之间调节推断过程。TwT分两个阶段训练:(1)对DeepSeek-R1提炼并经GPT-4o改写以体现类人推理经济性的难度感知长思维链轨迹进行监督微调;(2)采用混合奖励的强化学习优化翻译质量与推理效率。在覆盖域内、域外设置的15个基准,以及3种可见语言、59种未见语言上进行评估,通过对3种骨干模型的消融实验,TwT-7B和TwT-14B在翻译质量上优于大得多的现有最优(SOTA)推理模型,同时减少32%至60%的token使用量。这些结果证实,使翻译行为与认知原则对齐可实现MDMT中的强泛化、高翻译质量及高效推理。

英文摘要

Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators' ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resource-rational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-of-thought traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and out-of-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32--60\%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT.

Comments34 pages, 17 figures, and 21 tables. Accepted to ACL 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑