arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人类与LLM如何将性别解读进中性物理描述中

How Humans and LLMs Read Gender into "Gender-Neutral" Physical Descriptions

Yingjia Wan, Lin Lin, Elisa Kreiss

arXiv 2609.16366首次发表:更新:

发表机构

University of California, Los Angeles (UCLA)(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过GAPA数据集和16个LLM评估,首次实证表明中性物理描述在人类解读中仍保留系统性性别关联,且模型存在对齐偏差,挑战了性别中立沟通的假设。

AI 中文摘要

当基础模型描述人物时,近期关于AI公平性、可访问性和伦理的研究建议避免使用推断出的身份标签(如“她”、“他的”),转而采用看似“客观”的物理描述(如“短发”、“清晰的下颌线”)。然而,这类描述性语言是否能实现性别中立的沟通仍是一个悬而未决的实证问题。为研究此问题,我们引入了GAPA(物理属性的性别关联)数据集,该数据集包含来自不同来源的316个常见物理属性,并配有来自304位美国标注者的14,706条性别关联评分。结果显示,物理描述在读者中带有结构化和分级的性别关联,其中对女性和男性的关联比非二元身份更为一致和独特。接下来,我们评估了16个LLM,涵盖不同模型家族、规模和后训练变体,并与人类评分进行对比。这些模型部分恢复了人类的关联,但表现出系统性对齐偏差,包括评分分布压缩、对男性关联的对齐较弱,以及不成比例地针对非二元类别的不对称弃权(不执行)。最后,我们发布了表现最佳的代理模型,该模型训练用于预测人类对描述性语言的性别关联,并通过LitBank中人物描述的社会语言学分析展示了其实用性。综合来看,我们的发现首次提供了实证证据,表明看似“客观”的物理描述在人类解读中可能保留系统性的性别关联,并揭示了模型与人类错位的系统性模式。这对用物理描述替代显式性别标签必然产生性别中立沟通的假设提出了挑战,并强调了在人类-AI交互中使用此类描述传达主观身份类别的下游挑战。

英文摘要

When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., "she", "his") in favor of seemingly "objective" physical descriptions (e.g., "short hair", "a defined jawline"). Yet whether such descriptive language achieves gender-neutral communication remains an open empirical question. To study this, we introduce GAPA (Gender Associations of Physical Attributes), a dataset of 316 common physical attributes drawn from diverse sources, paired with 14,706 gender-association ratings from 304 US-based annotators. Results show that physical descriptions carry structured and graded gender associations among readers, with more consistent and distinctive associations for women and men than for non-binary identities. Next, we evaluate 16 LLMs across model families, sizes, and post-training variants against human ratings. The models partially recover human associations but exhibit systematic alignment biases, including compressed rating distributions, weaker alignment for associations with men, and asymmetric abstention that disproportionately targets the non-binary category. Finally, we release the best-performing proxy model trained to predict humans' gender associations of descriptive language and demonstrate its utility through a sociolinguistic analysis of character descriptions in LitBank. Together, our findings provide the first empirical evidence that seemingly "objective" physical descriptions can retain systematic gender associations in human interpretation, and uncover systematic patterns of model-human misalignment. This challenges the assumption that replacing explicit gender labels with physical descriptions necessarily yields gender-neutral communication, and highlights downstream challenges in using such descriptions to communicate subjective identity categories in human-AI interaction.

CommentsThe dataset and code are available at https://github.com/Yingjia-Wan/GAPA, and the predictor model is released at https://huggingface.co/alisa-yingjia-wan/gapa-predictor-olmo2-7b

Journal refIn Proceedings of Third Conference on Language Modeling (COLM), 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑