当大语言模型智能体进行谈判:供应链中的私人信息与动态议价
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
浏览论文内容
中文总结 AI 辅助
该研究以供应链议价问题为场景,将9个LLM与完美贝叶斯均衡基准测试,发现LLM智能体的谈判能力、提供商身份及提示策略对剩余分配和效率有显著影响,为AI智能体审计提供了三维框架。
中文摘要 AI 辅助
随着大语言模型(LLM)智能体从决策支持转向自主采购,企业需要了解委托的谈判代表是否创造价值、是否以可预测的方式分配价值,以及是否避免亏损合同。我们在一个典型的供应链议价问题中研究这一问题:拥有私人需求信息的买方与不知情的卖方就一份数量-支付合同进行谈判。我们在9840次LLM-LLM谈判中,将来自OpenAI、Google和阿里巴巴的9个LLM与经过验证的完美贝叶斯均衡进行基准测试。首先,能力决定价值创造:智能体在98.9%的谈判中达成一致,获得了95.4%未折现的最优剩余,但平均需要2.98轮谈判,而基准仅为1.25轮,这种延迟侵蚀了21%-34%的剩余;能力也决定可靠性:基线模型在19.2%的案例中接受个体非理性合同,而中端和旗舰模型的这一比例为0.0%-0.6%,这使得自动化利润验证成为低于该阈值的关键保障。其次,剩余分配具有关系性:提供商身份比能力排名更能预测谁能更好地获取剩余,自对弈买方份额平均为OpenAI的40%、Google的50%、阿里巴巴通义千问(Qwen)的70%,该排序在受限通信和无折现条件下依然存在;更换提供商卖方身份会使分配比例变动7-18个百分点,而强大的通义千问旗舰模型是跨系列卖方中最弱的,因此供应商选择是一阶分配决策。第三,提示是战略杠杆:委托将委托方的经济耐心与智能体的提示战略耐心分离开来,这一免费部署选择是剩余分配的单一最强驱动因素(占解释方差的90%)。这些共同确立了对AI智能体的均衡参考审计,涵盖三个维度:折现效率、分配特征和运营可靠性。
英文摘要
As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.
发表机构
- School of Business, University of Connecticut(康涅狄格大学商学院)
机构由 AI 辅助整理,请以论文原文为准。