arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估用于翻译错误检测的多语言句子嵌入:一项英-希腊语对比研究

Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study

Eleftherios Kalogeros, Athanasios Ntalakas, Manolis Gergatsoulis, Paschalis Nikolaou, Sotiria-Lito Alexaki

arXiv 2608.28776首次发表:更新:

发表机构

Ionian University; Laboratory on Digital Libraries and Electronic Publishing; Centre for Literary Translation and Book Policy(爱奥尼亚大学; 数字图书馆与电子出版实验室; 文学翻译与图书政策中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对比评估BGE-M3等5种多语言句子嵌入模型与COMETKiwi在英-希腊语翻译错误检测中的性能,发现嵌入模型与COMETKiwi的错误敏感性特征互补,嵌入模型更适合作为翻译评估框架的组成部分。

AI 中文摘要

多语言句子嵌入正越来越多地用于估计跨语言的语义相似度,但人们对其对细粒度翻译错误的敏感性仍了解不足。本研究探究通用多语言嵌入模型能否区分正确的英-希腊语翻译与经微小修改的错误替代译文。研究人员基于FLORES+句子对齐参考译文构建了对比数据集,并由两名翻译专家审核,该数据集包含1850个示例,涵盖10个核心错误类别和5个探索性错误类别,涉及事实、词汇语义、语法、关系、指称及语篇层面的现象。研究人员采用余弦相似度,对5种多语言句子嵌入模型(BGE-M3、Multilingual E5、Multilingual MPNet、LaBSE和Jina Embeddings v3)的英文源句与其正确及错误希腊语译文的相似度进行评估,同时也对无参考的COMETKiwi模型作为机器翻译质量评估基线进行评估。研究通过对比准确率和针对特定类别的敏感性得分差来评估性能,结果显示BGE-M3在嵌入模型中准确率最高,达89.30%,而COMETKiwi准确率为94.49%。嵌入模型检测明确的事实和词汇变化的可靠性高于时态-体和代词-指称错误;COMETKiwi在多个困难类别(包括时态和体、代词和指称、语义角色错误)上的性能有所提升,但对日期和时间错误的敏感性较低,且在数字检测上的表现逊于嵌入模型。研究结果表明,不同模型的错误敏感性特征具有互补性:多语言句子嵌入可提供有用的语义充分性信号,但更适合作为更广泛翻译评估框架的组成部分,而非独立指标。

英文摘要

Multilingual sentence embeddings are increasingly used to estimate semantic similarity across languages, yet their sensitivity to fine-grained translation errors remains insufficiently understood. This study investigates whether general-purpose multilingual embedding models can distinguish correct English-Greek translations from minimally modified erroneous alternatives. A contrastive dataset was developed from FLORES+ sentence-aligned reference translations and reviewed by two translation experts. It contains 1,850 examples across ten core and five exploratory error categories, covering factual, lexical-semantic, grammatical, relational, referential, and discourse-level phenomena. Five multilingual sentence-embedding models (BGE-M3, Multilingual E5, Multilingual MPNet, LaBSE, and Jina Embeddings v3) were evaluated using cosine similarity between each English source sentence and its correct and erroneous Greek translations. A reference-free COMETKiwi model was also evaluated as an MT quality-estimation baseline. Performance was assessed through contrastive accuracy and score margins for category-specific sensitivity. BGE-M3 achieved the highest accuracy among embedding models at 89.30 percent, while COMETKiwi achieved 94.49 percent. Embedding models detected explicit factual and lexical changes more reliably than tense-and-aspect and pronoun-coreference errors. COMETKiwi improved performance on several difficult categories, including tense and aspect, pronoun and coreference, and semantic-role errors, but showed lower sensitivity to date-and-time errors and underperformed the embedding models on numbers. The results show complementary error-sensitivity profiles: multilingual sentence embeddings provide useful semantic adequacy signals but are better suited as components of broader translation-evaluation frameworks than as standalone metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑