arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20925cs.CLcs.AI

无源机器翻译评估并非真正的机器翻译评估

Source-Free MT Evaluation Is Not MT Evaluation

Baban Gain, Ramakrishna Appicharla, Asif Ekbal

首次发表
浏览论文内容

中文总结 AI 辅助

本文指出无源机器翻译评估不符合翻译充分性定义,现有混合指标过度依赖参考,呼吁将质量估计(QE)作为基于源的充分性评估的主要方法,并设计优先考虑源-假设忠实度的混合指标。

中文摘要 AI 辅助

基于参考的指标仍是机器翻译评估的标准选择,部分原因是质量估计(QE)方法往往与人类判断的相关性较低。因此,无源、基于参考的评估已成为实际规范,尽管它不符合翻译充分性的定义,且对那些输出保留源文本含义但与参考不同的系统不公平。本文认为,充分性必须根据源文本进行判断,参考只是源文本的一种可能表述,可能引入偏差、欠规范或错误。我们进一步认为,源-参考-假设评估只有在评估者将参考视为辅助证据而非主要标准时才公平,否则,即使是感知源的评估也会将充分性简化为对参考的偏好。我们表明,与源相比,现有的混合指标高度依赖参考。我们的论点并非所有自动机器翻译指标都未使用源,而是任何移除源或允许参考主导源的评估协议,在充分性评估方面都存在结构性缺陷。然而,现有的机器翻译论文通常偏好基于参考的指标,仅在参考不可用时才使用QE指标。因此,我们呼吁将QE重新定义为基于源的充分性评估的主要方法,而非因参考缺失而采用的 fallback 方案;我们还呼吁设计混合指标时明确优先考虑源-假设的忠实度,仅将参考作为补充证据。

英文摘要

Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation has become the practical norm, even though it is unfaithful to the definition of translation adequacy and unfair to systems whose outputs preserve the source meaning while differing from the reference. This paper argues that adequacy must be judged with respect to the source. A reference is only one possible rendering of the source and may introduce bias, under-specification, or errors. We further argue that source-reference-hypothesis evaluation is fair only when the judge treats the reference as auxiliary evidence rather than as the primary standard. Otherwise, even source-aware evaluation can reduce adequacy to preference towards reference. We show the existing hybrid metrics are highly reliant on reference compared to source. Our argument is not that all automatic MT metrics fail to use the source. Rather, we argue that any evaluation protocol that removes the source, or allows the reference to dominate the source, is structurally incomplete for adequacy evaluation. However, existing MT papers generally prefer reference-based metrics and use QE metrics only when reference is unavailable. We therefore call for QE to be reframed as a primary approach to source-grounded adequacy evaluation, rather than as a fallback motivated by missing references. We further call for hybrid metrics whose designs explicitly prioritize source--hypothesis faithfulness while using references only as complementary evidence.

发表机构

  • Indian Institute of Technology Patna(印度巴特那印度理工学院)
  • Shiv Nadar University, Chennai(金奈希夫·纳达尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑