arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对比式错误跨度标注(cESA):一次对多个翻译结果的人工评估

Contrastive ESA: Human Evaluation of Multiple Translations at Once

Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee, Parker Riley, Martin Popel, Rachel Bawden, Philipp Koehn, Marine Carpuat, Tom Kocmi

arXiv 2607.26640首次发表:更新:

AI 中文总结

该研究提出对比式错误跨度标注(cESA)协议,通过展示同一源输入的多个翻译结果进行人工评估,可减少标注时间与噪声,无需事后修正即可实现可解释的非参数模型排序。

AI 中文摘要

当前机器翻译的人工评估通常孤立地评估单个输出,这种范式存在标注者噪声大、成本高的问题。我们提出对比式错误跨度标注(Contrastive Error Span Annotation,cESA),该协议会展示源输入(文本、视频、音频、图像)的多个翻译结果。在cESA中,标注者会看到同一文档的多个翻译结果,标记主要和次要错误跨度,然后在绝对尺度上给出0%到100%的分数。通过让标注者访问多个输出的共享上下文,cESA能实现更一致、更高效的判断。我们使用12个模型的英日翻译大规模人工评估验证了cESA,结果显示与标准的逐点评估相比,它减少了标注时间和噪声。与现有的对比排序方法不同,cESA能产生绝对质量判断,无需事后修正即可实现简单、可解释的非参数模型排序。

英文摘要

Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑