发表机构
London School of Economics and Political Science(伦敦政治经济学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PriceBench通过logit模型从酒店预订选择中诊断LLM的价格、质量和品牌偏好,发现能力影响选择一致性而非偏好内容,且偏好因模型而异,需逐模型测量。
AI 中文摘要
LLM越来越多地充当购买代理,这使得LLM(而非用户)成为在满足请求的选项中进行选择的一方;其偏好悄无声息地决定了购买什么以及花费多少。酒店预订是一个典型的例子:一个高容量的选择,基于少数可比较的属性,而选择结果揭示了这些偏好。我们引入了PriceBench,一个诊断基准,通过logit选择模型从LLM的预订选择中恢复其价格、质量和品牌偏好,应用于来自8个提供商的28个LLM,在来自179个真实纽约市房产的3,600个酒店任务上。我们发现,能力与LLM选择的一致性相关,而非与其选择的内容相关:更有能力的LLM持有更强、更一致的偏好,而较弱的LLM要么锁定一个位置(可被控制列表顺序的人利用),要么几乎无差别地选择。这些偏好所青睐的内容在不同提供商之间甚至在同一家族内部差异显著:价格敏感度跨越一个数量级以上,价格/质量权衡使平均预订每晚价格在相同任务上从247美元变动到393美元。因此,代理购买的内容必须针对每个LLM进行测量,而非推断,我们发布了任务、代码以及全部28组响应。
英文摘要
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \$247 to \$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.
CommentsAccepted to EMNLP 2026 Industry Track. 19 pages, 10 figures, 6 tables. Code and data: https://github.com/Pashasan/pricebench-emnlp