Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
智能体能够欺骗吗?使用社交推理游戏评估ParliamentBench中的推理与欺骗能力
机构 * University of Göttingen(哥廷根大学) ; National Institute of Informatics(信息学研究所) ; University of Tokyo(东京大学)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
AI总结 该研究基于《Secret Hitler》游戏构建ParliamentBench基准,评估16个LLM的欺骗与推理能力,发现前沿模型表现优异,多数LLM难以维持一致的欺骗人设。