大语言模型智能体能否进行竞争性定价?面向智能体商务的动态多属性拍卖基准
Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce
浏览论文内容
中文总结 AI 辅助
本研究推出动态多属性拍卖基准Bazaar,测试11个前沿LLM智能体的定价能力,发现其在客户获取与利润表现上存在差异,需求冲击下排名变化,最强智能体仅获不到三分之一最优利润,显示仍有提升空间。
中文摘要 AI 辅助
智能体商务正从概念走向已部署的基础设施:支付网络、零售商和AI平台正在为智能体代表商家和消费者进行交易搭建舞台。然而,这些智能体背后的大语言模型(LLM)是否能在真实市场中进行合理定价——在这类市场中,客户偏好是隐藏的、竞争对手会实时调整策略、需求可能毫无预兆地发生变化——这一点尚未得到系统测试。我们推出了Bazaar,这是一种适用于上述场景的动态密封投标多属性拍卖基准。尽管具有动态性,该基准仍基于闭式客户效用函数,可实现精确评估。在来自四家提供商的11个前沿大语言模型中,在客户获取方面领先的智能体(例如Gemini 3.1 Pro)通常并非在利润方面领先的智能体(例如Opus 4.6)。在需求冲击下,排名会再次发生变化:冲击前学习最快的智能体通常是之后修正信念最慢的,而Gemini 3.1 Pro尽管在利润方面不领先,但恢复速度最快。不过,即使是最强的智能体也只能获得不到事后最优利润的三分之一,这表明当前的大语言模型在智能体商务领域正在取得进展,但仍有很大的提升空间。
英文摘要
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.
发表机构
- Visa Research(维萨研究院)
机构由 AI 辅助整理,请以论文原文为准。