“如果我只能买一款:Galaxy S26 Ultra”——审计AI生成的产品推荐
"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations
浏览论文内容
中文总结 AI 辅助
本研究审计了主流AI聊天机器人的产品推荐,发现ChatGPT在79%的回复中表达第一人称偏好,且推荐结果和来源在不同条件下差异显著,强调独立审计需考虑重复响应和来源层级。
中文摘要 AI 辅助
消费者越来越多地使用AI聊天机器人获取购买建议。随着OpenAI和Google等公司通过广告将其AI商业化,这引发了关于此类建议的偏见和公正性的难题。为此,我们使用真实的商业咨询查询对流行的聊天机器人进行AI审计。首先,我们整理了一个包含2,528条真实商业咨询查询的数据集(ConsumerQ)。然后,我们评估了来自流行AI聊天机器人的1,536条产品查询回复:ChatGPT(聊天机器人和API)、Google Gemini(聊天机器人和API)以及Google Search(AI概览)。我们发现,ChatGPT在79%的产品推荐回复中表达了第一人称的产品偏好,而Gemini为7%,AI概览为2%,且推荐的产品在重复请求中经常变化。展示的来源差异很大:对于相同的查询,ChatGPT和Gemini界面平均仅共享5.4%的域名,在76.7%的比较中没有任何共同域名。API提供了与其对应界面不同的视角,ChatGPT的平均域名重叠率为12.0%,Gemini为14.8%,并且在暴露的来源信息的类型和层级上也存在差异。我们的研究结果表明,无论是孤立的回复还是API观察,都不能被视为消费者所遇到的商业建议的代表。因此,对AI中介的商业建议进行独立审计应考虑重复回复、面向消费者的条件以及所观察的来源层级。
英文摘要
Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this raises difficult questions about the bias and impartiality of such advice. In response, we conduct an AI audit of popular chatbots using real commercial-advice queries. First, we curate a dataset of 2,528 real commercial-advice queries (ConsumerQ). Then, we evaluate 1,536 responses to product queries from popular AI chatbots: ChatGPT (chatbot and API), Google Gemini (chatbot and API), and Google Search (AI Overviews). We find that ChatGPT expresses a first-person product preference in 79% of product-recommending responses, compared with 7% for Gemini and 2% for AI Overviews, while the products recommended often change across repeated requests. Displayed sources vary strongly: for the same query, the ChatGPT and Gemini interfaces share only 5.4% of domains on average, with no domain in common in 76.7% of comparisons. APIs provide a different view from their corresponding interfaces, with mean domain overlaps of 12.0% for ChatGPT and 14.8% for Gemini, and also differ in the types and layers of source information they expose. Our findings show that neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter. Independent audits of AI-mediated commercial advice should therefore account for repeated responses, consumer-facing conditions, and the source layer being observed.
发表机构
- Maastricht University(马斯特里赫特大学)
- Utrecht University(乌得勒支大学)
- University of Zurich(苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。