发表机构
National Yang Ming Chiao Tung University; Academia Sinica(国立阳明交通大学; 中央研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出狼人杀中的信念转变评估基准,通过分析 LLM 对指控消息的信念更新,发现大模型虽能更好区分狼人,但仍受指控影响,且难以整合指控内容与来源信任。
AI 中文摘要
诸如狼人杀之类的社交推理游戏越来越多地被用于评估 LLM 智能体,但现有的评估往往依赖于最终的游戏结果。我们提出了一个在狼人杀中的信念转变评估基准,通过信念更新来分析沟通技能。利用 LLM 游玩的游戏,我们标注了怀疑和指控消息,并测量一个观察的村庄侧模型的信念在每条消息后如何变化。我们在 1,224 条标注消息上评估了 40 种开放权重 LLM 配置。我们的结果显示,较大的模型能更好地根据游戏历史区分真正的狼人和村民,但指控仍然强烈影响它们的信念。模型对被指控的目标变得更加怀疑,而对指控者变得不那么怀疑,尤其是当指控者被信任时,即使指控者是狼人阵营。较大的模型能更好地抵抗来自它们已经不信任的指控者的指控。总体而言,我们的发现表明,当前高达 120B 参数的开放权重 LLM 在战略沟通中仍然难以将指控内容与来源信任整合。我们的基准和代码可在该 https URL 获取。
英文摘要
Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.
CommentsAccepted by the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026 Main Conference)