发表机构
Massachusetts Institute of Technology; University of Michigan, Ann Arbor; Shanghai Jiao Tong University; University of Hong Kong; McGill University; University of Chicago; Tongji University; Beijing Normal University; University of Texas at Dallas(麻省理工学院; 密歇根大学安娜堡分校; 上海交通大学; 香港大学; 麦吉尔大学; 芝加哥大学; 同济大学; 北京师范大学; 德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LLMAdBench基准,通过18000余条人类判断研究LLM回复中广告位置偏好,发现前沿LLM无法可靠替代人类评估,而微调的小模型可显著提升预测性能。
AI 中文摘要
将广告插入面向消费者的LLM输出中正成为一种新兴商业模式,但关于如何评估此类广告插入或其对用户偏好的影响,目前缺乏共享证据。我们引入LLMAdBench,一个用于研究LLM生成内容中广告投放的人类偏好基准。该基准聚焦于一个简单但具有实际重要性的决策:给定用户对话、LLM回复和匹配的广告,广告应放置在何处?我们的数据集比较仅广告位置不同、其他所有条件(包括用户查询、基础答案、广告和披露条件)均保持固定的回复对。人类标注者从广告主和用户两个视角,基于六项标准评估每一对回复。最终基准包含超过18000条人类判断,涵盖两种披露条件:明确将广告标记为赞助内容,以及将广告无披露地融入回复中。我们使用LLMAdBench评估八个前沿LLM作为偏好判断器,发现它们不能可靠地替代人类评估。即使最稳定的模型在展示顺序交换时也会反转约四分之一的决策,模型间一致性低,且其位置偏好与人类标注者存在系统性差异。此外,LLMAdBench包含大量可学习信号。特别是,在人类偏好上微调的Qwen3-8B模型相比其基础模型有显著提升,并在留出预测任务上优于所有零样本前沿判断器。除模型评估外,LLMAdBench还提供了关于广告主-用户权衡的定量证据,并表明赞助披露会系统性改变用户对广告位置的偏好。
英文摘要
Inserting advertisements (ads) into consumer-facing LLM output is emerging as a new business model, but there is little shared evidence on how such ad insertion should be evaluated or how it affects user preferences. We introduce LLMAdBench, a human-preference benchmark for studying advertising in LLM-generated content. The benchmark isolates a simple but practically important decision: given a user conversation, an LLM response, and a matched advertisement, where should the ad be placed? Our dataset compares pairs of responses that differ only in ad position while holding all other conditions fixed including the user query, base answer, advertisement, and disclosure condition. Human annotators evaluate each pair based on six criteria from both advertiser's and user's perspectives. The resulting benchmark contains more than 18000 human judgments across two disclosure conditions: explicitly labeling the ad as sponsored and merging it into the response without disclosure. We use LLMAdBench to evaluate eight frontier LLMs as preference judges and find that they are not reliable substitutes for human evaluation. Even the most stable models reverse roughly one quarter of their decisions when the presentation order is swapped, agreement across models is low, and their placement preferences differ systematically from those of human annotators. Moreover, LLMAdBench contains substantial learnable signal. In particular, a Qwen3-8B model fine-tuned on the human preferences improves substantially over its base model and outperforms all zero-shot frontier judges on the held-out prediction task. Beyond model evaluation, LLMAdBench provides quantitative evidence on the advertiser-user trade-off and shows that the sponsorship disclosure systematically changes users' preference over ad placement.