发表机构
Barcelona Supercomputer Center(巴塞罗那超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对术语感知翻译,提出仅保留模型翻译与术语表相矛盾的困难样本进行微调,将术语准确率从78.7%提升至89.9%,并构建SalamandraTA-7b系统,在WMT26评测中取得94.2%术语成功率。
AI 中文摘要
术语感知翻译要求的不只是正确的翻译:输出必须使用术语表规定的确切术语。标准的做法是在带有术语标注的翻译对上微调,但这种方法隐藏着一种低效性:对于大多数样本,术语表规定的正是模型本来就会生成的内容,因此这些样本对于学习遵循术语表毫无帮助。因此,我们只保留模型自身翻译与术语表相矛盾的样本。在固定数据量的受控研究中,仅这一筛选就将术语准确率从 78.7% 提升到了 89.9%。筛选后的数据由基于开放模型的双向合成流程构建,是我们公开发布的 SalamandraTA-7b-instruct v3.0 指令微调混合数据的一部分。该模型按原样使用,并封装在文档级推理流程中,构成了 BSC 对 WMT26 术语共享任务赛道 1 的提交。在官方 WMT26 评测中,我们的系统在 74.6 chrF++ 下达到了 94.2% 的术语成功率,22 个提交中仅有 2 个在两个指标上均优于我们。在去年的基准上,它也超越了基于 GRPO 的系统,尽管它仅使用普通的监督微调进行训练。
英文摘要
Terminology-aware translation asks for more than a correct translation: the output must use the exact terms a glossary prescribes. The standard recipe, fine-tuning on glossary-annotated translation pairs, hides an inefficiency: for most examples the glossary prescribes exactly what the model would have produced anyway, so they teach nothing about following a glossary. We therefore keep only the examples where the model's own translation contradicts the glossary. In a controlled study at fixed data volume, this selection alone raises term accuracy from 78.7% to 89.9%. The filtered data, built by a two-way synthetic pipeline on open models, is part of the instruction-tuning mixture of our public release SalamandraTA-7b-instruct v3.0, which, used exactly as released and wrapped in a document-level inference pipeline, forms the BSC submission to the WMT26 Terminology Shared Task Track 1. At the official WMT26 evaluation, our system achieves 94.2% term success at 74.6 chrF++, with only two of the twenty-two submissions outperforming it on both metrics. On last year's benchmark, it also surpasses our GRPO-based system, despite being trained solely with ordinary supervised fine-tuning.
CommentsTo appear at Proceedings of the Eleventh Conference on Machine Translation (WMT26; camera-ready version)