发表机构
Inria; Sorbonne Université; CNRS; ISIR(法国国家信息与自动化研究所; 索邦大学; 法国国家科学研究中心; 智能系统与机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器翻译术语评估忽视人类译者变体的问题,提出变体感知评估方法,结合术语表准确性、一致性和跨术语变体度量,发现MT变体少于人类且术语表约束抑制有效变体。
AI 中文摘要
机器翻译(MT)中的术语评估通常假设每个源术语只有一个正确的目标形式。然而,人类译者通常会引入变体,而当前指标将其视为不一致并予以惩罚。我们探讨了如何在英法科学翻译的文档级MT评估中考虑这种变体,结合了基于术语表的准确性、翻译一致性,以及一种新的跨术语变体(CTV)诊断度量,该度量测试变体关系是否在语言间得到保留。基于对两个平行语料库的分析(这些语料库由四个MT系统翻译),我们发现:(1)MT系统生成的目标端变体少于人类译者;(2)迁移模式强烈依赖于变体类型;(3)一致性排名随度量选择而变化;(4)使用术语表约束MT可提高准确性和一致性,但通过抑制有效变体而降低了CTV。我们主张采用变体感知的评估方法,即根据目标端变体是否反映源端变体来调节一致性惩罚。
英文摘要
Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source term. However, human translators routinely introduce variation that current metrics penalize as inconsistency. We examine how to account for this variation in document-level MT evaluation of English-French scientific translation, combining glossary-based accuracy, translation consistency, and a new cross-term variation (CTV) diagnostic measure that tests whether variation relationships are preserved across languages. Based on analyses of two parallel corpora, translated by four MT systems, we find that (1) MT systems generate less target-side variation than human translators; (2) transfer patterns strongly depend on the variation type; (3) consistency rankings vary with the choice of metric; and (4) constraining MT with a glossary improves accuracy and consistency but degrades CTV by suppressing valid variation. We argue for variation-aware evaluation that conditions consistency penalties on whether target-side variation mirrors source-side variation.
Journal refAMTA 2026 - 17th Conference of the Association for Machine Translation in the Americas, Aug 2026, Quebec City, Canada