SocialGrid:一种用于具身多智能体系统规划与社交推理的基准测试
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
- Technical University of Darmstadt(达姆施塔特技术大学)
- German Research Center for Artificial Intelligence(德国人工智能研究中心)
- Lab1141(Lab1141实验室)
- Centre for Cognitive Science, Darmstadt(达姆施塔特认知科学中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SocialGrid通过评估LLM在具身多智能体环境中的规划、任务执行与社交推理能力,揭示了现有模型在任务完成和社交推理上的不足,并提供规划 oracle 和竞争排行榜以促进改进。
AI中文摘要:
随着大型语言模型(LLMs)从文本处理器转变为自主代理,评估其在具身多代理设置中的社交推理能力变得至关重要。我们引入了SocialGrid,一个受Among Us启发的具身多代理环境,用于评估LLM代理在规划、任务执行和社交推理方面的表现。我们的评估表明,即使是最强大的开源模型(GPT-OSS-120B)在任务完成和规划方面的准确率也低于60%,代理会陷入重复行为或无法导航基本障碍。由于糟糕的导航会干扰社交智能的评估,SocialGrid提供了一个可选的规划 oracle 来将社交推理与规划缺陷分离。虽然规划辅助提高了任务完成度,但社交推理仍然是瓶颈:无论规模如何,代理在检测欺骗时都仅靠浅层启发式方法,而非积累行为证据。SocialGrid提供了自动失败分析和细粒度指标,使开发者能够诊断并改进其代理。我们还通过Elo评分建立了竞争排行榜,利用对抗性联赛比赛来促进改进。
英文摘要:
As Large Language Models (LLMs) transition from text processors to autonomous agents, evaluating their social reasoning in embodied multi-agent settings becomes critical. We introduce SocialGrid, an embodied multi-agent environment inspired by Among Us that evaluates LLM agents on planning, task execution, and social reasoning. Our evaluations reveal that even the strongest open model (GPT-OSS-120B) achieves below 60% accuracy in task completion and planning, with agents getting stuck in repetitive behaviors or failing to navigate basic obstacles. Since poor navigation confounds evaluation of social intelligence, SocialGrid offers an optional Planning Oracle to isolate social reasoning from planning deficits. While planning assistance improves task completion, social reasoning remains a bottleneck: agents fail to detect deception at near-random chance regardless of scale, relying on shallow heuristics rather than accumulating behavioral evidence. SocialGrid provides automatic failure analysis and fine-grained metrics, enabling developers to diagnose and improve their agents. We also establish a competitive leaderboard using Elo ratings from adversarial league play.