发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨密集检索器对查询中身份信号(政治倾向和方言)的敏感性,发现检索结果存在偏见,可能加剧极化与健康差距。
AI 中文摘要
密集检索器决定了哪些文档能够到达用户以及使用这些文档的语言模型,然而它们通常是在中性查询下进行评估的。我们探究真实用户在查询中表达的身份信号——政治意识形态和方言——是否会使检索器返回的结果产生偏差。我们在两个领域设计了评估:政治新闻和消费者健康问题,每个领域都将仅改变身份信号的受控合成集与自然主义查询配对。在五个密集检索器和一个稀疏基线中,每个检索器(i)都会检索与查询自身政治倾向相符的文章,并且(ii)在针对非裔美国人语言(AAL)书写的问题上的表现,比针对白人主流英语(WME)的问题更差。两项分析将这些差距与查询的身份信号联系起来,而不仅仅是表面词汇:剔除一个聚合的词汇不对称得分后,合成差距在很大程度上仍然存在;线性探针能够从检索器的查询嵌入中恢复出倾向和方言,而这些信息超出了词元级别的特征。如果不加以解决,这种检索偏差可能会加剧极化,并强化AAL使用者已经面临的健康差距。代码可在以下网址获取:https://this https URL。
英文摘要
Dense retrievers decide which documents reach users and the language models that use them, yet they are typically evaluated with neutral queries. We ask whether the identity signals that real users express in their queries---political ideology and dialect---bias what a retriever returns. We design evaluations in two domains, political news and consumer-health questions, each pairing a controlled synthetic set that varies only the identity signal with naturalistic queries. Across five dense retrievers and a sparse baseline, every retriever (i) retrieves articles that align with the query's own political lean and (ii) performs worse for questions written in African American Language (AAL) than in White Mainstream English (WME). Two analyses tie these gaps to queries' identity signals beyond surface vocabulary: partialling out an aggregate lexical-asymmetry score leaves the synthetic gaps largely intact, and linear probes recover lean and dialect from the retrievers' query embeddings beyond token-level features. Left unaddressed, such retrieval biases risk contributing to polarization and reinforcing the health disparities already faced by AAL speakers. Code is available at https://github.com/Andrewtcr/bias-ret.
CommentsEMNLP 2026 camera-ready, with a correction to Fig. 4