arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18618cs.CLcs.AIcs.LG

LatentMT:具有潜在推理的机器翻译

LatentMT: Machine Translation with Latent Reasoning

Wei-Rui Chen, Samar M. Magdy, Chiyu Zhang, Wenhui Zhu, Zhipeng Wang, Muhammad Abdul-Mageed

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出LatentMT,将潜在推理循环语言模型用于机器翻译,采用小型主干模型经轻量级训练,在多语言翻译方向表现出色,循环推理早期提升质量后期饱和,且训练和推理计算效率高,为紧凑高效的机器翻译提供新路径。

中文摘要 AI 辅助

潜在推理循环语言模型(LoopLMs)为机器翻译(MT)提供了不同的扩展路径:它们在隐藏状态内进行额外的循环计算,而非增加参数数量或发出明确的思维链令牌。我们引入LatentMT,这是对用于机器翻译的潜在推理LoopLMs的首次系统研究。LatentMT采用小型26亿参数主干模型并进行轻量级训练。在涵盖高、中、低资源语言的32个翻译方向上,LatentMT取得了与三到五倍大的模型相当的性能。在高资源语言中具有竞争力,在中低资源语言上达到了当前最优性能。研究循环推理步骤数量的扩展行为发现,早期循环计算持续提高翻译质量,随后迅速饱和。机理分析表明隐藏表示差异沿循环推理步骤轴缩小,支持了性能饱和现象。效率分析表明LatentMT在训练和推理计算方面比具有相似性能的非潜在推理大模型更低,使潜在循环计算成为实现紧凑、高效且强大的机器翻译的有前途路径。

英文摘要

Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT): instead of increasing parameter count or emitting explicit chain-of-thought tokens, they spend additional recurrent computation inside hidden states. We introduce LatentMT, the first systematic study of latent-reasoning LoopLMs for machine translation. LatentMT adapts a small 2.6B-parameter backbone model with lightweight training. Across 32 translation directions spanning high-, mid-, and low-resource languages, LatentMT achieves performance comparable to models three to five times larger. It is competitive in a high-resource language and achieves state-of-the-art performance on both mid-resource and low-resource languages. Studying the behavior of scaling the number of recurrent reasoning steps, we find that recurrent computation consistently improves translation quality in early steps, then saturates quickly afterwards. Our mechanistic analysis shows that hidden-representation differences shrink along the recurrent reasoning-step axis, supporting the observed saturation in performance. Finally, our efficiency analysis shows that LatentMT requires lower training and inference compute than much larger non-latent-reasoning models with similar performance, making latent recurrent computation a promising path toward compact, efficient, and strong machine translation.

↑