arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13916cs.CL

North Small Translate:先进的高性价比翻译(Cohere CAT+)

North Small Translate: Advanced Cost-Effective Translation (Cohere CAT+)

  • Cohere

机构由 AI 辅助整理,请以论文原文为准。

Tom Kocmi, Alexandre Bérard, Phil Blunsom, Samuel Cahyawijaya, Shaun Cassini, Nicholas Frosst, Ona de Gibert, Aidan Gomez, Nithya Govindarajan, Shun Kiyono, Oli… 展开作者

Tom Kocmi, Alexandre Bérard, Phil Blunsom, Samuel Cahyawijaya, Shaun Cassini, Nicholas Frosst, Ona de Gibert, Aidan Gomez, Nithya Govindarajan, Shun Kiyono, Olivia Lasche, Lawrence Rogers, Kelly Marchisio, Nikita Moghe, Yash More, Camila Moran-Hidalgo, Yiyang Nan, Michael Sachs, Trisha Starostina, Daan van Stigt, Spencer Rarrick, Sebastian Vincent, Ivan Zhang

AI总结:

本文提出 North Small Translate,一个基于混合专家架构的高性价比翻译模型,通过难度采样和五步训练协议,在 50 种语言上达到顶尖翻译性能,无需推理开销。

AI中文摘要:

我们推出 North Small Translate,这是一个基于 LLM 的开源权重机器翻译(MT)模型,具备指令跟随能力,其基础与 Cohere 的 Command A Plus 相同,采用混合专家架构,总参数为 2180 亿,其中 250 亿为活跃参数。North Small Translate 使用难度采样来获取具有挑战性的文档,并通过五步训练协议进行训练,该协议结合了监督微调、直接偏好优化和在线强化学习。我们通过非推理基础模型优先考虑吞吐量,并辅以可选的智能体能力以解锁翻译质量的提升。North Small Translate 经过训练可执行与 MT 相关的任务,包括译后编辑和质量估计,以及通用指令跟随等相关任务。该模型在参数低于 1 万亿的模型类别中,跨越 50 种语言取得了顶尖的 MT 性能,且无需在推理时运行昂贵的推理过程。

英文摘要:

We present North Small Translate, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same foundation as Cohere's Command A Plus, a mixture-of-experts architecture with 25 billion active parameters out of 218 billion total parameters. North Small Translate is trained using difficulty sampling to obtain challenging documents and a five-step training protocol combining supervised fine-tuning, direct preference optimization, and online reinforcement learning. We prioritized throughput through a non-reasoning base model and supplemented with optional agentic capabilities to unlock translation quality gains. North Small Translate is trained to perform MT-related tasks, including post-editing and quality estimation, as well as related tasks such as general instruction following. The model achieves top MT performance across 50 languages in the class of models under 1T parameters, with no need to run expensive reasoning at inference time.

↑