发表机构
Instituto de Telecomunicações; National Technical University of Athens; University of Amsterdam(电信研究所; 雅典国立技术大学; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在WMT 2026共享任务中,基于GAMBIT+职业平衡子集,评估七个语言对的机器翻译评估指标,发现阳性译文普遍得分更高且存在职业刻板差异,表明性别偏见仍存,需多维分析。
AI 中文摘要
性别偏见在机器翻译(MT)中仍然是一个持续存在的问题,它不仅影响生成的译文,也影响其自动评估。当源文本未指明人物性别时,译文可能以阳性或阴性形式呈现该人物,而尽管源文本未提供任何依据,机器翻译系统和评估指标都可能在这些形式之间表现出系统性偏好。我们在WMT 2026自动翻译质量评估系统共享任务中,使用GAMBIT+中一个职业平衡的子集来研究这一行为。我们考虑了七个英语源语言对,其中六个来自原始数据集,目标语言为阿拉伯语、捷克语、希腊语、冰岛语、俄语和乌克兰语,并将原始资源扩展至德语。该子集包含每种目标语言1,308对阳性/阴性译文对,涵盖436个ISCO-08职业组中的每个职业三个示例。我们评估了共享任务提交系统和基线在分数预测和错误标注方面的表现,考察了性别相关差异的方向、幅度和频率。我们发现总体上阳性译文获得更高分数的趋势,以及不同职业间遵循刻板性别表征的差异,尽管这种偏好的强度和一致性在不同评估者和语言间差异显著。我们的结果表明,性别偏见在机器翻译评估中仍然存在,但要捕捉其程度,需要超越单一的聚合指标,转向评估者行为的互补维度。
英文摘要
Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.
CommentsAccepted for publication at the 11th Conference of Machine Translation (WMT26), co-located with EMNLP 2026