在错误的路灯下探索:自动翻译质量评估的局限性
Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation
浏览论文内容
中文总结 AI 辅助
本文从理论和实证角度分析自动翻译质量评估(QE)的结构性局限,指出其无法作为独立工具用于实际翻译工作流,建议未来聚焦于基于MQM的人类评估自动化。
中文摘要 AI 辅助
翻译质量评估(QE)的自动化已成为大规模管理翻译质量的广泛讨论方法,为实现该目标,越来越多的工具和技术被发布。然而,新QE系统的激增并未总是伴随着稳健、透明且可复现的研究与测试,这一缺口值得批判性审视。本文从理论和实证角度考察QE技术的一些基本局限性,认为当前QE系统在结构上无法作为可靠的独立工具应用于实际翻译工作流。所考察的证据表明,QE存在一系列相互关联且大多未解决的局限性:最根本的是,对孤立片段层面翻译质量的评估存在问题,因为它往往会忽略衔接、连贯以及文体和修辞文本特征;此外,实证研究还记录了其他若干局限性和缺陷,包括泛化失败、系统性偏差、过拟合与分布崩溃、性能差距、错误标注挑战以及数据稀缺。这些是由人类语言和翻译作为认知与交际行为的复杂性所导致的结构性局限性,截至目前,更多数据和更好的架构尚未能克服这些局限性。因此,片段层面的QE分数不应在生产环境中用作路由、发布或绕过审核的独立依据;我们认为未来工作应聚焦于基于MQM(多维度质量度量)实现人类评估的自动化。
英文摘要
Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing number of tools and technologies have been released in pursuit of this goal. However, the proliferation of new QE systems has not always been accompanied by robust, transparent, and reproducible research and testing. This gap deserves critical scrutiny. This paper examines some fundamental limitations of the QE technology from both theoretical and empirical perspectives, arguing that current QE systems are structurally ill-equipped to serve as reliable standalone tools in real-world translation workflows. The reviewed evidence suggests that QE suffers from a range of interrelated and largely unresolved limitations. Most fundamentally, the evaluation of the quality of translation at the level of isolated segments is problematic because it tends to miss out on cohesion, coherence, and stylistic and rhetorical text features. In addition, empirical research documents several other limitations and flaws, including failure to generalize, systematic biases, overfitting and distribution collapse, performance gaps, error annotation challenges, and data scarcity. These are structural limitations arising from the complexity of human language and translation as a cognitive and communicative act - limitations that more data and better architectures have so far not overcome. Consequently, segment-level QE scores should not be used as a standalone basis for routing, release, or review bypass in production; we argue future work should focus on automating human evaluation grounded in MQM.