arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使COMET在不同文字系统间可比:印度语机器翻译评估中分词器引发的文字偏差的诊断与修正

Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation

G. L. John Salvin, Swapnil Hingmire

arXiv 2610.08159首次发表:更新:

发表机构

Indian Institute of Technology Palakkad(印度理工学院帕拉卡德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对COMET在跨文字系统比较中的偏差,提出COMET-QN归一化方法,消除文字间评分范围差异,将标注一致性从0.300提升至0.399,并保留语言内排序。

AI 中文摘要

COMET将翻译质量报告为单一数字,且该数字通常被用于比较以不同文字书写的目标语言。这种比较隐含了文字不变性假设:评分不应依赖于承载目标的书写系统。我们通过在IndicMT Eval上将目标重新编码为拉丁文字来测试该假设,此举在保持内容和人工评分不变的同时改变了正字法形式。结果显示,文字身份可解释原生文字COMET方差的22.9%,且在全部五种研究语言中,与标注者的一致性均有所下降。我们将该效应追溯到分词器,并使用三种无标签诊断进行测量。该偏差包含两个缺陷,而非一个。来自不同文字的评分占据不兼容的范围,且在单一文字内,该指标对翻译排序的准确性降低。没有任何保序变换能修复第二个缺陷。第一个缺陷可被COMET-QN精确消除,该方法将每个(语言,文字)对的评分分布映射到共享参考上。与标注者的合并一致性从0.300提升至0.399,这使得来自不同文字的评分可安全地置于同一坐标轴上,且每个语言内的排序均被证明得以保留。基于奇偶特征的回归器进一步恢复了17.1%的损失灵敏度。剩余部分属于编码器,任何后处理都无法触及。因此,我们建议发布归一化评分、三种诊断及其计算所依据的分词器身份,以便读者能辨别评分中有多少反映翻译质量,多少反映书写系统。

英文摘要

COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target into Latin script, which changes orthographic form while holding content and human ratings fixed. Script identity then accounts for 22.9% of native-script COMET variance, and agreement with annotators falls in all five languages studied. We trace the effect to the tokeniser and measure it with three label-free diagnostics. The bias is two faults, not one. Scores from different scripts occupy incompatible ranges, and within a single script the metric orders translations less accurately. No order-preserving transform of the score can repair the second fault. The first is removed exactly by COMET-QN, which maps the score distribution of each (language, script) pair onto a shared reference. Pooled agreement with annotators rises from 0.300 to 0.399, which is what makes scores from different scripts safe to place on one axis, and every within-language ordering is provably preserved. A regressor over parity features recovers a further 17.1% of the lost sensitivity. The remainder belongs to the encoder, and no post-processing can reach it. We therefore recommend publishing the normalised score, the three diagnostics, and the identity of the tokeniser they were computed against, so that a reader can tell how much of a score reflects translation quality and how much reflects the writing system.

Comments18 pages, 2 figures. Camera-ready version, accepted at WMT 2026. Code and data: https://github.com/John-salvin/script-bias-comet-normalisation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑