arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

持久伙伴提高学习智能体的定价

Persistent Partners Raise Prices Among Learning Agents

Paul-Peter Arslan, Yubin Kim, Xiao Xiao

arXiv 2609.35402首次发表:更新:

发表机构

Institute For Future Technologies; Devinci Higher Education; Massachusetts Institute of Technology(未来技术研究院; 达芬奇高等教育学院; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究平台配对如何影响学习智能体的定价,发现保持同一伙伴显著提高利润水平,且该效应在隐藏对手价格时依然存在,但并非源于习得性惩罚。

AI 中文摘要

当定价智能体在平台上反复相遇时,平台决定谁与谁配对。我们探究这一选择是否改变智能体学习到的价格,以及价格上升是否伴随习得性惩罚。在Calvano等人的Bertrand双寡头模型的预注册随机实验中,每个智能体的价格由表格型Q学习模块设定,而非其所附带的小型语言模型,我们随机化每个智能体是否保持其伙伴、是否看到对手价格以及能否发送消息。保持同一伙伴可使训练平均利润水平提高竞争利润与垄断利润之间差距的0.27(95%置信区间0.20至0.35,所有二十对配对运行均为正),这是我们的注册主要结果,并使静止价格提高Nash到垄断范围的0.17(事后分析)。一个纯表格学习器在全部25个额外区块中重现了该效应,且在那里一个永久伙伴比约三个伙伴更能提高水平(+0.23对比+0.05,探索性)。当对手价格隐藏时,定价模块无法看到降价,因此无法惩罚,但静止价格上升幅度相同且持续到训练结束,而对手可见时,随着训练时间延长,上升幅度缩小(事后分析)。当对手可见时,静态最优反应者解释了强制偏离探针所读取的惩罚的三分之一到一半,在对手能看到降价的起始状态上,扣除该效应后,习得性惩罚的注册测试结果不确定。仅寻找惩罚的测试会遗漏对手隐藏时的价格上升,而检查有利可图的偏离则会标记其中大多数价格(事后分析)。在探索性扩展中,未训练的Qwen2.5 7B和14B模型在单一提示下,当对手价格从提示中省略时表现出该效应,而当对手价格显示时则表现不一致,7B结果在新区块上重复,而另外两个模型家族未表现出该效应。

英文摘要

When pricing agents meet repeatedly on a platform, the platform decides who faces whom. We ask whether that choice moves the prices the agents learn, and whether a rise comes with learned punishment. In a pre-registered randomised experiment in the Bertrand duopoly of Calvano et al., each agent's price is set by a tabular Q-learning module, not by the small language model attached to it, and we randomise whether each agent keeps its partner, sees its rival's prices and can send messages. Keeping the same partner raises the level of profits, averaged over training, by 0.27 of the gap between competitive and monopoly profit (95% CI 0.20 to 0.35, all twenty paired runs positive), our registered primary result, and the resting price by 0.17 of the Nash-to-monopoly range (post hoc). A plain tabular learner reproduces the effect in all 25 further blocks, and there one permanent partner raises the level more than about three do (+0.23 against +0.05, exploratory). Where rival prices are hidden, the price-setting module cannot see a cut, so cannot punish it, yet the resting price rises as much and the rise lasts to the end of training, while with visible rivals it shrinks with longer training (post hoc). Where the rival is visible, a static best responder accounts for a third to a half of what a forced-deviation probe reads as punishment, on the starts where the rival can see the cut, and net of it the registered test of learned punishment is inconclusive. A test that looks only for punishment would thus miss the rise where the rival is hidden, while a check for profitable deviations flags most of those prices (post hoc). In an exploratory extension, untrained Qwen2.5 7B and 14B models under one prompt show the effect when the rival's price is left out of the prompt and inconsistently when it is shown, the 7B result replicating on fresh blocks, while two other model families show none.

Comments29 pages, 5 figures. Pre-registered on OSF (https://osf.io/98bx5, under embargo). Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑