arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12125cs.GTcs.AIcs.CLcs.MA

大语言模型会照顾同类吗?相似性信号可诱导合作

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer

中文总结 AI 辅助

该研究提出首个带分级相似性信号的LLM决策评估框架,发现不同LLM应对相似性信号差异大、数据集对诱导合作影响小等,构建的LLM-行为博弈论模型可在高相似性下支持合作均衡。

中文摘要 AI 辅助

随着带有用户指定目标的基于大语言模型(LLM)的智能体被广泛部署,它们在策略性交互中相遇的情况越来越多,面临着找到互利结果的挑战。现有文献认为,在智能体知晓自身遵循高度相似决策模式的场景中,诸如囚徒困境之类的合作问题是可解决的,例如在单一文化的AI生态系统中。受此研究方向启发,本文引入了首个用于评估LLM决策的框架,该框架中智能体被提供分级相似性信号。我们的研究发现包括:不同的LLM模型在应对相似性信号时表现出巨大差异,部分现代模型在合作问题、收益结构和提示框架中展现出一致行为;出人意料的是,用于计算相似性信号的数据集对诱导合作的影响很小甚至没有,且当被要求自行评估另一模型的思维链推理时,LLM模型会系统性地自我认定为高度相似;最后,我们开发了一个LLM-行为博弈论模型,该模型捕捉了它们的部分推理逻辑,且表明在足够高的相似性得分下,该模型可在均衡状态下支持合作结果。

英文摘要

As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.

补充信息

↑