arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同博弈,不同故事:大语言模型智能体的最小保守战略稳健性基准

Same Game, Different Story: A Minimal Conservative Strategic Robustness Benchmark for Large Language Model Agents

Seyed Pouyan Mousavi Davoudi, Arshia Gharagozlou, Alireza Amiri-Margavi, Amin Gholami Davodi, Hamidreza Hasani Balyani

arXiv 2607.19670首次发表:更新:

AI 中文总结

研究大语言模型智能体在战略环境中的可靠性,通过“相同博弈,不同故事”基准,以收益不变框架变化下行动分布不变性定义稳健性,经二次分析已发表数据得出社会关系框架会改变模型行为,应分别评估稳健性与能力。

AI 中文摘要

大语言模型智能体越来越多地在战略环境中运行,结果取决于其他智能体的行动。这就引发了一个可靠性问题:当相同的激励通过不同的叙述呈现时,模型会做出一致的选择吗?我们引入了“相同博弈,不同故事”基准,将战略稳健性定义为在收益保持不变的框架变化下,模型诱导行动分布的不变性。我们通过对已发表的GPT - 3.5、GPT - 4和LLaMa - 2在四个社会困境博弈中的总体合作率进行二次分析来说明该框架。由于无法获得试验级数据,从已发表的数据中重建了近似计数。结果表明,即使基础行动集和收益保持不变,社会关系框架也能显著改变大语言模型的行为。因此,应使用收益等效提示族而非单一博弈呈现来分别评估战略稳健性和战略能力。

英文摘要

Large language model agents are increasingly deployed in settings where the value of an action depends on what other agents do. This creates a strategic reliability problem: the same game may be described as a business negotiation, a friendly compromise, a diplomatic exchange, or an abstract payoff matrix, and the model may choose different actions even when the incentives are unchanged. This paper introduces \emph{Same Game, Different Story}, a benchmark for strategic robustness: invariance of model-induced action distributions under payoff-preserving language changes. The empirical analysis uses a deliberately narrow, literature-calibrated comparison from Lorè and Heydari's peer-reviewed study: business framing versus friend-sharing framing across GPT-3.5, GPT-4, and LLaMa-2 in four social-dilemma games, with 300 initializations per retained model-game-context cell. The retained design comprises 24 of the source study's 60 cells, representing 7,200 decisions. Because trial-level files were not available from the article, the analysis is presented as a secondary calibration based on reconstructed published rates, not as new model runs. As a conservative sensitivity analysis, effect magnitudes are attenuated by 30\% toward the null: action shifts are multiplied by 0.70, and non-robustness, defined as one minus the robustness score, is multiplied by 0.70. Under this attenuation, pooled strategic robustness is 0.783 with a 95\% bootstrap interval from 0.774 to 0.790, and friend-sharing framing raises cooperation by 0.307 with a 95\% bootstrap interval from 0.297 to 0.316 relative to business framing. The analysis supports the narrower claim that social-relational framing can change strategic choices even when incentives are held fixed, without extending the analysis to a broader suite of contextual or cross-benchmark comparisons.

Comments7 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑