arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习销售:多产品市场中战略性大型语言模型智能体的强化学习

Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets

Shuze Daniel Liu, Claire Chen, Jiuqi Wang, David Simchi-Levi, Thorsten Joachims

arXiv 2609.33289首次发表:更新:

AI 中文总结

本研究提出一种基于可验证奖励强化学习的后训练方法,使大型语言模型智能体在多产品讨价还价环境中作为卖方动态匹配买家与产品,在剩余提取和分配质量上达到或超越万亿参数模型,并泛化至未见市场。

AI 中文摘要

在多产品市场中运行的自主大型语言模型(LLM)智能体必须在信息不对称和资源约束下做出顺序决策。我们开发了一种机器学习方法,用于训练此类智能体在多物品讨价还价环境中作为卖方有效行动,其中卖方同时与一组独立买家就一系列可替代资产进行谈判。买家对产品持有私有的、异质的估值,并且每个买家最多只能购买一件物品。面对总通信轮次的限制,卖方必须根据买家的私有估值动态地将买家与最有利可图的产品进行匹配,同时战略性地将其有限的交互预算分配给具有更大潜在价值的组合。我们将此问题形式化为部分可观察马尔可夫决策过程,使用结构化的四部分消息协议,将自然语言映射到可解析且受约束的决策空间。利用这种形式化,我们设计了一种使用可验证奖励强化学习(RLVR)的后训练方法。为了评估该框架,我们构建了一个多维指标套件,用于量化约束遵守、卖方剩余提取和分配质量。我们训练的卖方智能体学会更有效地将有限库存与买家匹配,在卖方剩余提取和买家-产品分配质量方面均达到或超越万亿参数前沿模型。最后,这些学习到的策略对未见过的市场结构、相关估值分布和训练期间未遇到的价格范围具有稳健的泛化能力。

英文摘要

Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, where a seller concurrently negotiates a catalog of substitutable assets across a pool of independent buyers. Buyers hold private, heterogeneous valuations across products, and each can purchase at most one item. Facing limits on total communication turns, the seller must dynamically match buyers with the most profitable products considering their private valuations, while strategically allocating its limited interaction budget toward combinations of greater potential value. We formalize this problem as a Partially Observable Markov Decision Process using a structured, four-part message protocol that maps natural language into a parsable and regulated decision space. Using this formalization, we design a post-training method using Reinforcement Learning from Verifiable Rewards (RLVR). To evaluate this framework, we construct a multidimensional metric suite that quantifies constraint adherence, seller surplus extraction, and allocation quality. Our trained seller agent learns to match limited inventory to buyers more effectively, matching or outperforming trillion-parameter frontier models in both seller surplus extraction and buyer-product allocation quality. Finally, these learned strategies generalize robustly to unseen market structures, correlated valuation distributions, and price ranges not encountered during training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑