arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05949cs.CLcs.HC

文档级翻译评估的盲区

The Blindness of Document-Level Translation Evaluation

Ahrii Kim, Vilém Zouhar, Chanjun Park, Seong-heum Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过反事实混合段落实验,发现文档级翻译评估的得分、排名和错误标注在连贯与不连贯文档间无显著差异,表明评估协议存在盲区,未能有效捕捉文档级质量。

中文摘要 AI 辅助

文档级机器翻译(MT)评估通过向标注者展示完整文档来扩展段落级协议,其假设是这种展示方式能够引发文档级别的判断。我们通过一个反事实条件(MIX)来检验这一假设,在该条件下,每个文档由来自不同系统的段落混合而成,从而在保留文档级展示的同时打破跨段一致性。在18,420条专家级英语到韩语标注和14个自动评估指标中,连贯文档与不连贯文档之间的得分、系统排名和错误标注在统计上等效。感知并不能解释这一现象:在展示匹配的段落时,评估者在87.3%的试验中能够识别出连贯的段落是单一译者的作品。文档展示确实改变了标注者的工作方式,但这种改变并未反映在记录的输出中。存在盲区的是评估协议,而非标注者。问题不在于得分偏低,而在于投入在文档级系统、指标和标注上的资源可能并未衡量其预期衡量的内容。

英文摘要

Document-level machine translation (MT) evaluation extends segment-level protocols by presenting full documents to annotators, on the assumption that such presentation elicits document-level judgments. We test this assumption with a counterfactual condition (MIX) in which each document combines segments drawn from different systems, preserving document-level presentation while breaking cross-segment consistency. Across 18,420 expert Englis-to-Korean annotations and 14 automatic metrics, scores, system rankings, and error annotations are statistically equivalent between coherent and incoherent documents. Perception does not explain this: shown matched passages, raters identify the coherent one as the work of a single translator in 87.3% of trials. Document presentation does change how annotators work, but that change does not reach the recorded output. What is blind is the protocol, not the annotator. The concern is not that scores fall short, but that the resources invested in document-level systems, metrics, and annotation may not be measuring what they are intended to measure.

发表机构

  • AI-Bio Convergence Research Inst.(AI-生物融合研究所)
  • ETH Zurich(苏黎世联邦理工学院)
  • Soongsil University(崇实大学)
  • School of Software(软件学院)
  • Dept. of Intelligent Semiconductors(智能半导体系)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑