发表机构
Marburg University(马尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文构建德语-英语多语言基准数据集,评估8种语言模型的反LGBTQ偏见,发现模型存在相关刻板印象,微调可平均降偏但效果不均,凸显多语言偏见评估需文化适应。
AI 中文摘要
尽管语言模型中的性别与种族偏见已被广泛研究,但反LGBTQ偏见仍未得到充分探索,尤其是在英语之外的语言中。现有基准往往无法捕捉文化与语言差异,且依赖性别表征。本文引入了一个用于评估语言模型中反LGBTQ偏见的德语-英语多语言基准数据集,它结合了德语区酷儿群体提供的刻板印象,以及WinoQueer的德语译本。该数据被用于评估8个不同规模和架构的语言模型,并探索通过在社区及进步媒体内容上进行微调来缓解偏见。结果显示,语言模型会复制反酷儿刻板印象,且在不同身份与模型间存在差异。翻译数据与社区来源数据的差异凸显了文化适应对多语言偏见评估的重要性。微调平均而言可降低偏见,但在不同模型与身份间的效果并不一致。警告:本文包含反酷儿仇恨言论及刻板印象的示例。
英文摘要
While gender and racial biases in language models have been widely studied, anti-LGBTQ biases remain underexplored, particularly beyond English. Existing benchmarks often do not capture cultural and linguistic variation and rely on gender representations. This paper introduces a multilingual German-English benchmark dataset for the evaluation of anti-LGBTQ biases in language models. It combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer. The data is used to evaluate eight language models across sizes and architectures and explore mitigation through fine-tuning on community and progressive media content. Results show that language models reproduce anti-queer stereotypes, with variation across identities and models. Differences between the translated and community-based data highlight the importance of cultural adaptation for multilingual bias evaluation. Fine-tuning reduces bias on average, but not consistently across models and identities. Warning: This text contains examples of anti-queer hateful language and stereotypes.
CommentsAccepted at EMNLP 2026