发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出Social Gym多智能体社会游戏环境与SPaRTan自博弈反思迁移方法,前者可客观基准测试LLM社会推理,后者可提升GPT-5-mini的弱角色表现,为改进LLM社会推理提供了可复现基础。
AI 中文摘要
LLM智能体越来越多地被部署在多智能体社会场景中,必须进行合作、协商并适应其他智能体。衡量和改进这些社会技能十分困难,因为与数学或逻辑不同,社会互动没有客观的真实值:评估依赖于LLM评判,这一方式成本高昂、主观且存在噪声,模型无法获得可靠的学习信号。为解决这两个问题,我们首先推出Social Gym,这是一个包含21种多智能体社会游戏(例如狼人杀、抵抗组织、间谍牌)的环境,其由规则决定的结果使智能体的表现可验证且客观,搭配Elo锦标赛生成跨游戏排行榜。基准测试实验显示,尽管GPT-5-mini位居排行榜榜首,但没有任何模型能在所有游戏或所有游戏角色中均表现出色,这凸显了社会推理的局限性。受此启发,我们还提出了SPaRTan(自博弈与反思迁移),这是一种无需训练的自我改进循环:模型玩游戏、反思其轨迹及结果以生成可迁移的剧本,并在后续游戏中应用该剧本。我们的结果表明,SPaRTan剧本可帮助GPT-5-mini智能体提升其较弱角色的表现,但在很大程度上无法改进Qwen3-32B的表现。综上,Social Gym与SPaRTan为无需权重更新即可衡量和改进LLM社会推理提供了可复现、可验证的基础。
英文摘要
LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, we first introduce Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, we additionally propose SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Our results show that SPaRTan playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B's performance. Together, Social Gym and SPaRTan offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.