LocQE:利用后期编辑实现本地化质量评估的原则性域适应
LocQE: Principled Domain Adaptation for Localisation Quality Estimation by Leveraging Post-Edits
- Technical University of Munich(慕尼黑工业大学)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
- LILT
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对质量评估模型在本地化领域域适应差的问题,提出利用后期编辑数据通过多任务微调和分词器干预,显著提升区分优选与拒绝翻译的能力。
AI中文摘要:
学习型质量评估(QE)模型(如COMETKiwi)广泛使用,且在通用机器翻译评估中表现良好。然而,这些模型在未见过的领域上表现不佳,限制了它们在真实世界本地化场景中的性能。我们表明,这些模型对本地化中的一些重要因素不敏感,例如数字是否被准确翻译,甚至翻译中是否保留了正确的空格数量和标点符号。此外,机器翻译优化的一个关键能力是QE模型能够准确地对单个片段的不同翻译进行排序,而这种能力在域迁移中受到显著影响。在缺乏大规模直接评估数据的情况下,我们提出了原则性的微调方法,即使使用少量的后期编辑数据也能缩小领域差距。通过采用多任务微调方法和简单的分词器干预,我们创建了一个QE模型,该模型在本地化场景中区分首选后期编辑与拒绝的初始翻译方面明显更优。我们表明,偏好和人工连续分数相互稳定,并认为为了校准指标(无论是绝对分数还是同一源文本不同翻译之间的比较),这两种类型的信号都是必需的。
英文摘要:
Learned quality estimation (QE) models such as COMETKiwi are widespread and work well for general machine translation evaluation. However, they are known to struggle on unseen domains, limiting their performance in a real-world localisation context. We show that they are insensitive to some important factors in localisation, such as whether numbers are translated accurately, or even whether the correct number of spaces and punctuation are preserved in a translation. Further, a key capability for optimisation of machine translation is the ability of QE models to accurately rank different translations of a single segment, which suffers significantly from the domain transfer. In the absence of large-scale direct assessment data, we propose principled fine-tuning approaches to reduce the domain gap with even small amounts of post-editing data. Using a multi-task fine-tuning approach and a simple tokeniser intervention, we create a QE model which proves markedly better at distinguishing preferred post-edits from rejected initial translations in a localisation context. We show that preferences and artificial continuous scores stabilise each other, and argue that to calibrate metrics both in terms of their absolute scores and comparisons between translation of the same source, both types of signal are needed.