arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解释GAND:关于性别模糊自然数据与对比归因的资源

Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution

Janiça Hackenbuchner, Jasper Degraeuwe, Arda Tezcan, Joke Daems

arXiv 2607.22546首次发表:更新:

发表机构

Ghent University(根特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究机器翻译中性别偏见问题,提出GAND基准资源,通过将其部分内容翻译成两种语言并扩展对比翻译,经特征归因分析揭示影响目标翻译中模糊实体性别翻译的源词。

AI 中文摘要

机器翻译(MT)系统持续产生有性别偏见的翻译。在自我表达至关重要的时代,基于默认行为和刻板印象的误译会对这些系统的用户造成伤害。为更好理解这些系统在缺乏明确性别线索时如何翻译性别,我们需要能自然反映性别模糊场景的基准资源。为此,我们展示了GAND,一个用于MT的性别模糊自然数据基准资源,由英语源句子组成,专门设计用于分析语境线索对翻译中性别影响。我们利用GAND进行可解释性分析:将GAND的一个子集翻译成两种语法性别的语言,并通过人工制作的对比翻译进行扩展。随后的特征归因分析揭示了上下文中影响目标翻译中模糊指称实体性别翻译的源词。

英文摘要

Machine translation (MT) systems continue to produce gender-biased translations. In a time where self-expression is paramount, mistranslations based on default behaviour and stereotyping can lead to harm for users of these systems. To better understand how these systems translate gender in the absence of clear gender cues, we need benchmarking resources that reflect gender-ambiguous scenarios in a natural way. To this end, we present GAND, a gender-ambiguous natural data benchmarking resource for MT consisting of English source sentences, specifically designed to analyse the influence of contextual cues on gender in translation. We leverage GAND to conduct an interpretability analysis: we translate a subset of GAND into two grammatical gender languages and extend these with manually crafted contrastive translations. A following feature attribution analysis reveals source words in context that inform the gender translation of an ambiguous referent entity in the target translation.

CommentsAccepted at EAMT2026: Technical Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑