arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型在广告相关性判断中的偏见系统性研究

A Systematic Investigation of Bias in Large Language Models for Advertising Relevance

Weiwei Wang, Yinchuan Xu, Jialu Gao, Youkow Homma, Jian Jiao

arXiv 2610.07544首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统调查了LLM在广告相关性判断中的偏见,发现广告主身份、语言和人口统计措辞影响评估结果,并探讨了推理和训练阶段的缓解措施及其有效性条件。

AI 中文摘要

大型语言模型(LLMs)越来越多地被用于判断广告与查询的匹配程度,但这些判断的公平性却很少受到关注。我们对LLMs在查询和广告相关性判断中的公平性进行了系统性研究。我们的反事实框架考察了广告主身份、可能的流行度、输入语言以及人口统计措辞的影响。我们研究了作为分类相关性评判者的GPT-4o以及一个专门为相关性预测训练的Qwen-7B模型。广告主和语言实验使用了从真实广告日志中采样的查询和广告对。受控的合成查询用于研究就业、住房和信贷中的人口统计关联。对于这两个模型,改变广告主身份或输入语言都可能改变相关性评估。选定的人口统计比较也显示出与常见刻板印象一致的模式,特别是涉及性别和职业的模式。我们进一步研究了模型推理和训练期间的缓解措施。结果表明,其有效性取决于广告主信息是否与查询相关,以及广告主标签在训练数据中的分布方式。这些发现可以帮助广告从业者识别公平性风险,并为LLM相关性系统制定合适的缓解方法。

英文摘要

Large language models (LLMs) are increasingly used to judge how well an advertisement matches a query, but the fairness of these judgments has received limited attention. We conduct a systematic study of fairness in relevance judgments made by LLMs for queries and advertisements. Our counterfactual framework examines the effects of advertiser identity and possible popularity, input language, and demographic wording. We study GPT-4o as a categorical relevance judge and a Qwen-7B model trained specifically for relevance prediction. The advertiser and language experiments use query and advertisement pairs sampled from real advertising logs. Controlled synthetic queries are used to study demographic associations in employment, housing, and credit. For both models, changing the advertiser identity or input language can alter the relevance assessment. Selected demographic comparisons also show patterns consistent with common stereotypes, particularly those involving gender and occupation. We further study mitigation during model inference and training. The results indicate that its effectiveness depends on whether advertiser information is relevant to the query and how advertiser labels are distributed in the training data. These findings can help advertising practitioners identify fairness risks and develop suitable mitigation methods for LLM relevance systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑