arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型智能体中基于声誉的合作的涌现

Emergence of Reputation-Based Cooperation in LLM Agents

Kazuya Horibe, Kenji Itao, Wataru Toyokawa

arXiv 2608.04507首次发表:更新:

AI 中文总结

研究发现LLM智能体的合作鲁棒性由对手禀赋敏感性(图像评分机制)和叛逃者排除严格程度决定,其无法发展更鲁棒的Leading-Eight规范,凸显文化进化的LLM合作的漏洞并提出自下而上规范构建的需求。

AI 中文摘要

大型语言模型(LLM)智能体之间的合作能否在搭便车者入侵下保持进化稳定性?我们研究了一种间接互惠捐赠博弈,其中LLM智能体观察行为轨迹并按连续尺度进行捐赠。策略以自然语言提示的形式呈现,通过跨代的文化传播进化。在四种LLM后端中,对搭便车者入侵的鲁棒性差异超过一个数量级。这种鲁棒性的最强预测因子是对手禀赋敏感性,即智能体区分合作与不合作对手的程度,该指标将经典的图像评分机制(Image Scoring)进行了操作化。相比之下,对Leading-Eight L1范数的遵守程度并不能预测鲁棒性。鲁棒性取决于叛逃者排除机制:虽然不同模型中合作者奖励和叛逃者惩罚存在差异,但仅叛逃者排除的严格程度可预测对搭便车者入侵的抵抗能力。这些发现表明,LLM智能体被限制在类似图像评分的区分机制中,无法发展出更鲁棒的Leading-Eight规范,凸显了文化进化的LLM合作中的根本性漏洞,并为自下而上的规范构建方法提供了动机。

英文摘要

Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion? We study an indirect reciprocity donation game where LLM agents observe behavioral traces and donate on a continuous scale. Strategies, represented as natural language prompts, evolve through cultural transmission across generations. Across four LLM backends, robustness to free-rider invasion varies by more than an order of magnitude. The strongest predictor of this robustness is opponent endowment sensitivity, the degree to which agents discriminate between cooperative and uncooperative opponents, operationalizing the classical Image Scoring mechanism. By contrast, adherence to the Leading-Eight L1 norm does not predict robustness. Robustness depends on defector exclusion: while both cooperator reward and defector punishment vary across models, only the stringency of defector exclusion predicts resistance to free-rider invasion. These findings reveal that LLM agents are confined to Image Scoring-like discrimination and fail to develop the more robust Leading-Eight norms, highlighting a fundamental vulnerability in culturally evolved LLM cooperation and motivating bottom-up approaches to norm construction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑