发表机构
Inria; Sorbonne Université; CNRS; ISIR(法国国家信息与自动化研究所; 索邦大学; 法国国家科学研究中心; 智能系统与机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TermJudge是一种文档级术语评估指标,采用两步LLM-as-judge程序区分术语错误与有效变异,在系统级和段落级元评估中排名第一,并证明术语表注入能通过消除真实错误来改善翻译。
AI 中文摘要
现有的用于评估机器翻译(MT)中术语使用的自动指标会对任何偏离固定参考译法的行为进行惩罚,从而将翻译错误与人类译员通常会产生的有效术语变异混为一谈。我们引入了TermJudge,一种文档级术语指标,它为每个术语出现位置赋予一个可解释的判定:符合术语表的出现被确定性地判定,而偏离则在两步LLM-as-judge程序下,利用完整文档上下文进行评估:第一步检测并标记术语错误;第二步将有效的文档级变异与不一致之处区分开来。经过专家错误标注和文档级人工MQM评分的验证,TermJudge在系统级和段落级元评估中均排名第一,领先于术语表一致性和质量估计基线。当应用于八种系统翻译学术文档时,在两种提示条件下,我们观察到术语表注入在所有配对比较中均改善了术语翻译,这是通过消除真正的错误而非有效的变异来实现的。TermJudge以开源代码形式发布。
英文摘要
Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminology metric that assigns an interpretable verdict to every term occurrence: glossary-conforming occurrences are settled deterministically, while divergences are assessed under a two-step LLM-as-judge procedure using the full document context: the first detects and labels terminology errors; the second sorts valid document-level variations from inconsistencies. Validated against expert error annotations and document-level human MQM scores, TermJudge ranks first in both system- and segment-level meta-evaluation, ahead of glossary-conformity and quality-estimation baselines. When applied to eight systems translating academic documents, under two prompting conditions, we observe that glossary injection improves terminology translation in all paired comparisons, by removing genuine errors rather than valid variation. TermJudge is released as open-source code.
Journal refWMT 2026 - The Eleventh Conference in Machine Translation 2026, Oct 2026, Budapest, Hungary