在盲区中:为机器翻译评估构建伪参考
In the Blind: Building Pseudo-References for MT Evaluation
浏览论文内容
中文总结 AI 辅助
针对无人工参考的语言对,提出用多模型多提示生成候选、无参考QE评分及按文档选择器构建伪参考,并加入语言识别惩罚解决错误语言排序问题,校准于WMT25人工判断,发布方法及来源。
中文摘要 AI 辅助
WMT26通用机器翻译任务在10个没有人工参考(既非从零翻译,也非人工对机器翻译输出进行后编辑)的语言对上评估系统。我们描述了如何为这些语言对以及其他六个具有某种形式人工参考的语言对构建伪参考:七个模型在最多五种提示条件下翻译了3,277份官方文档,总共产生26种系统-提示组合;然后三个无参考质量估计(QE)模型对每个候选进行评分;一个按文档选择器挑选一个翻译,GPT-5.5在需要时进行后编辑。在没有参考的情况下工作暴露了QE引导选择的一个失败模式:度量标准将错误语言中的流畅输出排在正确翻译之上。在评分融合中加入置信度缩放的语言识别惩罚将错误语言计数降至零,并且由此产生的选择器在MetricX上的得分仍然优于它所取代的排名融合基线。由于在构建这些语言对的伪参考时没有可用参考,我们根据去年的WMT25人工判断来校准每个选择决策。构建后发布的人工评估显示了选择错误的代价:当选择器保留了前沿模型候选时,我们的参考与最强参赛系统持平;当未保留时,则比它们低最多17个ESA点。我们发布了选择方法和每个参考的来源(此https URL)。
英文摘要
The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these pairs and six other language pairs (in which some forms of human references are available): seven models translate the 3,277 official documents under up to five prompt conditions, giving a total of 26 system-prompt combinations; then three reference-free quality estimation (QE) models score every candidate; and a per-document selector picks one translation, which GPT-5.5 post-edits where needed. Working without references exposed a failure mode of QE-guided selection: the metrics rank fluent output in the wrong language above correct translations. Adding a confidence-scaled language identification penalty to the score fusion drives the wrong-language count to zero, and the resulting selector still scores better on MetricX than the rank-fusion baseline it replaces. Since no references were available for these pairs while we were building them, we calibrate every selection decision on last year's WMT25 human judgments. The human evaluation, released after construction, shows the cost of getting selection wrong: our references stand with the strongest participating systems when the selector kept a frontier-model candidate, and fall up to 17 ESA points below them when it did not. We release the selection method and the provenance of every reference (https://github.com/surrey-nlp/PseudoRef)
发表机构
- University of Surrey(萨里大学)
- National Research Council Canada(加拿大国家研究委员会)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。