发表机构
DEVCOM Army Research Laboratory; Gifu University; University of Southern California(DEVCOM陆军研究实验室; 岐阜大学; 南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过迭代囚徒困境博弈比较推理型与非推理型Gen AI模型,发现声誉、策略和情感影响合作,推理型模型更依赖策略与声誉,并表现出异质终局行为,强调需标准化合作基准以指导负责任部署。
AI 中文摘要
随着生成式人工智能(Gen AI)系统在经济和社会性重要互动中承担越来越自主的角色,理解它们的合作倾向——以及塑造这种倾向的信号——已变得至关重要。我们使用迭代囚徒困境博弈考察了前沿Gen AI模型的合作行为,操纵了对手声誉(正面、未知、负面)、策略(剥削性vs.慷慨性)以及非语言情感信号(传达竞争性或合作性评价的面部表情)。在第一项针对非推理型模型(Claude 3.5、Gemini 2.0 Flash、GPT-4o)的研究中,合作行为系统地受到所有三个因素的影响,与人类行为研究中长期记录的模式相似,尽管各模型对每个因素的权重存在显著差异。第二项针对推理型模型(Claude 4.6、Gemini 3、GPT-5.2)的研究揭示了更集中于策略和声誉的依赖,几乎消除了非推理型模型中观察到的“波将金效应”(在诊断性和谐博弈中表现为近乎一致的合作),以及情感在符合层级线索加权策略而非简单丧失社会敏感性的条件下发挥的作用。推理型模型还表现出异质的终局行为,从持续合作到系统性最后一轮背叛效应,揭示了特定模型的可利用性特征,这对谈判及其他多轮互动中的部署具有直接的实际意义。综合来看,这些发现将Gen AI模型刻画为日益复杂但异质的社会行动者,并强调了开发标准化合作基准的实际价值,以指导Gen AI在互动性、社会性重要场景中的负责任部署。
英文摘要
As generative AI (Gen AI) systems take on increasingly autonomous roles in economically and socially consequential interactions, understanding their propensity to cooperate -- and the signals that shape this propensity -- has become essential. We examine cooperative behavior in frontier Gen AI models using the iterated prisoner's dilemma, manipulating counterpart reputation (positive, unknown, negative), strategy (extortion vs. generosity), and non-verbal emotional signaling (facial expressions conveying competitive or cooperative appraisals). In a first study with non-reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT-4o), cooperation was systematically shaped by all three factors, paralleling patterns long documented in human behavioral research, though models varied substantially in how heavily each factor was weighted. A second study with reasoning models (Claude 4.6, Gemini 3, GPT-5.2) revealed a more concentrated reliance on strategy and reputation, a near-elimination of the Potemkin effect observed in non-reasoning models (evidenced by near-uniform cooperation in a diagnostic harmony game), and a more conditional role for emotion consistent with a hierarchical cue-weighting strategy rather than a simple loss of social sensitivity. Reasoning models also showed heterogeneous end-game behavior, ranging from sustained cooperation to systematic last-round defection effect, revealing model-specific exploitability profiles with direct practical relevance for deployment in negotiation and other multi-round interactions. Together, these findings characterize Gen AI models as increasingly sophisticated, though heterogeneous, social actors, and underscore the practical value of developing standardized cooperation benchmarks to inform the responsible deployment of Gen AI in interactive, socially consequential settings.