arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体能够欺骗吗?使用社交推理游戏评估ParliamentBench中的推理与欺骗能力

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa, Jan Philip Wahle, Bela Gipp, Terry Ruas

arXiv 2607.28146首次发表:更新:

发表机构

University of Göttingen; National Institute of Informatics; University of Tokyo(哥廷根大学; 信息学研究所; 东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究基于《Secret Hitler》游戏构建ParliamentBench基准,评估16个LLM的欺骗与推理能力,发现前沿模型表现优异,多数LLM难以维持一致的欺骗人设。

AI 中文摘要

当大语言模型(LLM)作为智能体部署在医疗、法律等高风险场景中时,理解其欺骗能力对安全至关重要。受控的社交推理游戏提供了可复现的代理,用于隔离和评估这些复杂的对抗性行为。我们提出基于《Secret Hitler》游戏的开源基准框架ParliamentBench,用于在需要欺骗、说服和信息不对称下推理的场景中评估LLM。我们在1600场模拟对局中评估16个LLM,包括它们之间对弈、与人类对弈,并与大量在线对局进行比较。我们引入三个新指标,分别用于衡量社交推理、推理和欺骗一致性。实验表明,前沿模型在合作角色和欺骗角色中均表现出色,形成了由GPT-5.4、Kimi K2.5、Grok 4.1 Fast和DeepSeek 3.1 Terminus组成的前四强集群,而最弱的模型表现低于随机(33%)和简单算法(45%)基线。大多数LLM难以在整场游戏中保持一致的欺骗人设,欺骗保留率降至50%以下。

英文摘要

As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑