arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于机器翻译质量标注的大语言模型:人类与模型均面临挑战

Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged

Hala Almaghout, Christian Federmann, Qin Gao

arXiv 2610.10918首次发表:更新:

发表机构

Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文评估LLMs在MQM和ESA两种MT质量标注方案上与人类标注者的一致性,发现二者在多数场景下不可靠,明确了人类与LLM各自的标注挑战,提出可通过人类-LLM协作标注流程提升可靠性。

AI 中文摘要

大语言模型(LLMs)被视为机器翻译(MT)评估中人类判断的更高效、更具成本效益的替代方案。由于机器翻译评估涵盖大量语言对、领域及标注粒度,LLMs在被可靠用作人类评估的替代方案之前,必须在这些维度上接受全面评估。本文通过将LLMs与人类标注者的一致性进行比较,评估了LLMs在两种主流机器翻译质量评估方案——多维质量度量(MQM)和错误跨度标注(ESA)上的性能。我们在包含70种语言对的长上下文测试集以及公开可用的WMT23和WMT25数据上呈现了结果,调查了不同语言对和领域下的评分与错误跨度标注一致性。我们的结果表明,尽管在部分评估任务中LLMs与人类标注者的一致性超过人类标注者之间的一致性,但两者在不同标注方案、语言对和领域间差异显著,且在大多数场景下仍不可靠。此外,我们明确了人类和LLM标注者面临的挑战:人类在细粒度MQM标注和低资源语言对上尤其面临挑战,而LLMs则在 minor errors( minor errors 保留原文)、错误语言变体及错误跨度标注上存在困难。我们的结果凸显了人类和LLM标注性能均有改进潜力,可能通过人类-LLM协作标注流程解决本研究中发现的可靠性问题。

英文摘要

Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation. With MT evaluation spanning a large number of language pairs, domains and levels of annotation granularity, LLMs must be thoroughly evaluated across these dimensions before being reliably used as alternatives to human evaluation. In this paper, we evaluate the performance of LLMs for two prominent MT quality evaluation schemes: Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) by comparing their agreement with human annotators. We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains. Our results show that while LLM agreement with human annotators exceeds agreement between human annotators for some evaluation tasks, both vary substantially across annotation schemes, language pairs and domains and remain unreliable for most settings. Furthermore, we identify challenges facing both human and LLM annotators: humans are particularly challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggle with minor errors, wrong language variants and error span annotation. Our results highlight the potential for improvement for both human and LLM annotation performance, possibly through human-LLM collaborative annotation pipelines that address the reliability issues identified in this work.

CommentsAccepted at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑