廉价交谈稳定LLM智能体中的策略互动
Cheap Talk Stabilizes Strategic Interaction in LLM Agents
中文总结 AI 辅助
本研究探讨廉价交谈对LLM智能体重复博弈中策略持续性的影响,发现其能稳定互动轨迹,且效果与模型及历史相关。
中文摘要 AI 辅助
大型语言模型日益被部署为交互式智能体,这使得它们的行动策略在重复交互中的持续性对于可靠的多智能体操作至关重要。我们研究了智能体生成的、非约束性的预交流(“廉价交谈”)是否以及如何提高这种持续性,实验涉及四个开放权重的7-9B参数LLM。我们的实验涵盖四个重复的双人博弈——囚徒困境、雪堆博弈、猎鹿博弈和和谐博弈——其激励结构从战略冲突到协调一致不等,每个博弈在六种情境下呈现。我们在所有四个博弈中观察到不稳定的轨迹,尽管其普遍性和程度强烈依赖于模型和情境。跨模型、博弈和情境,廉价交谈主要起稳定作用,五次修正的逆转集中在社会或团队框架中;效果因模型和情境而异。受控的当前消息干预在Qwen中识别出两个可分离的输出级通道:降低行动不确定性和减少轮间行动概率漂移。匹配的历史-消息反事实进一步表明,最近伙伴行为调节了互惠利益语言与自我优先语言对策略持续性的影响。最后,在囚徒困境中,我们在Qwen和Falcon的晚期Transformer层中识别出一个历史平衡的策略-内容方向;投影出该方向会增加闭环游戏中的实际切换,表明完整轨迹对该组件具有因果敏感性。总之,这些发现表明,廉价交谈可以使个体轨迹在多样化的激励结构中更加持续,同时揭示稳定的幅度和机制依赖于模型和历史。
英文摘要
Large language models are increasingly deployed as interacting agents, making the persistence of their action policies across repeated interaction critical for reliable multi-agent operation. We investigate whether and how agent-generated, non-binding pre-play communication ("cheap talk") increases such persistence in four open-weight 7-9B-parameter LLMs. Our experiments span four repeated two-player games -- Prisoner's Dilemma, Snowdrift, Stag Hunt, and Harmony -- with incentive structures ranging from strategic conflict to alignment, each presented in six contexts. We observe unstable trajectories in all four games, although their prevalence and magnitude depend strongly on model and context. Across models, games, and contexts, cheap talk is predominantly stabilizing, with five corrected reversals concentrated in social or team framings; effects vary substantially by model and context. Controlled current-message interventions identify two separable output-level channels in Qwen: reduced action uncertainty and less between-round drift in action probabilities. Matched history-by-message counterfactuals further show that recent partner behavior conditions how mutual-benefit versus self-prioritizing language affects policy persistence. Finally, in Prisoner's Dilemma, we identify in Qwen and Falcon a history-balanced policy-content direction in late transformer layers; projecting out this direction increases realized switching during closed-loop play, demonstrating that complete trajectories are causally sensitive to this component. Together, these findings show that cheap talk can make individual trajectories more persistent across diverse incentive structures, while revealing that the magnitude and mechanisms of stabilization are model- and history-dependent.