发表机构
Northwestern University; Boston University(西北大学; 波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一个评估LLM品牌推荐的框架,通过BRP@k和MRR@k指标及重复采样,发现推荐显著性更受市场可见性影响,且品牌可条件性检索,强调应视推荐为随机检索排名过程。
AI 中文摘要
大语言模型(LLMs)越来越多地被用于产品推荐,但评估其推荐效果所面临的挑战不同于传统信息检索和推荐系统。LLMs可以在没有显式候选集的情况下生成推荐,并且对同一查询的重复响应可能产生不同的品牌和排名。我们引入了一个评估开放式LLM品牌推荐的框架,该框架独立于模型输出定义竞争集,并通过重复采样估计推荐的普遍性和显著性。我们使用品牌推荐概率(BRP@$k$)和平均倒数排名(MRR@$k$)来操作化这些概念,并将该框架应用于五个产品类别中的六个LLMs。仅类别查询揭示了知名品牌的大量遗漏,且推荐显著性遵循传统品牌知名度的证据有限。相反,显著性与更广泛的市场可见性信号相关,特别是搜索兴趣和在线品牌讨论。基于需求的查询表明,对用户目标和约束进行情境化会改变检索到的品牌,而诊断性定位探针表明,当提供独特线索时,从普通推荐中遗漏的品牌仍可条件性地被检索到。这些发现强调,应将LLM推荐视为一种随机检索与排名过程来评估,而非基于单个生成的列表。我们提供开源软件和数据,以支持对LLM生成的品牌推荐进行可复现的评估。
英文摘要
Large language models (LLMs) are increasingly used for product recommendation, but evaluating their recommendations presents challenges that differ from conventional information retrieval and recommender systems. LLMs can generate recommendations without an explicit candidate set, and repeated responses to the same query can produce different brands and rankings. We introduce a framework for evaluating open-ended LLM brand recommendations that defines the competitive set independently of model outputs and estimates recommendation prevalence and prominence through repeated sampling. We operationalize these constructs using Brand Recommendation Probability (BRP@$k$) and Mean Reciprocal Rank (MRR@$k$), and apply the framework to six LLMs across five product categories. Category-only queries reveal substantial omission of established brands and limited evidence that recommendation prominence follows conventional brand popularity. Instead, prominence is associated with broader marketplace-visibility signals, particularly search interest and online brand conversation. Needs-based queries show that contextualizing users' goals and constraints changes which brands are retrieved, while diagnostic positioning probes demonstrate that brands omitted from ordinary recommendations can remain conditionally retrievable when distinctive cues are supplied. These findings highlight the need to evaluate LLM recommendation as a stochastic retrieval-and-ranking process rather than from individual generated lists. We provide open-source software and data to support reproducible evaluation of LLM-generated brand recommendations.