arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在隐藏信息社交推理游戏中审计基于信念的大语言模型智能体

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games

Yuan Gao, Jiangyi Yang, Yao Zhao, Yichi Zhang

arXiv 2607.10814首次发表:更新:

发表机构

University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在9人狼人杀环境中评估大语言模型智能体的难题,构建可审计框架,通过1080场游戏实验发现主动信念条件下好人方结果更好,虽未明确信念内容作用机制,但框架有使效应可测等贡献,将外部信念定位为可审计认知基线。

AI 中文摘要

在隐藏信息的多智能体环境中评估大语言模型智能体具有挑战性,最终结果具有高方差,且很少能揭示智能体决策的原因。我们在9人狼人杀环境中对此进行研究,智能体在严格的代码级信息隔离下行动。我们构建了一个可审计框架,该框架维护关于隐藏角色的外部信念状态,记录信念更新和信念 - 行动偏差作为结构化证据,并支持防御性离线改进循环。在1080场冻结游戏中,包括信念禁用、主动信念、内核消融、阵营限制、消耗策略和高负载分支,主动信念条件与更好的好人方结果相关。在200个种子的A0/A1比较中,好人方胜率从0.205提高到0.390(配对McNemar卡方检验\(\chi^2 = 16.4\),\(p < 0.001\)),不可逆的女巫毒人错误减少。然而,我们并未将这种转变归因于信念内容。直接的行动 - 信念一致性较低(约为0.21),仅将信念给予狼人对好人方的帮助大于仅给予好人方,这反对简单的持有者利益解释。我们将此效应报告为一种关联,并将其机制视为未解决。贡献在于审计框架本身:它使效应可测量,揭示低直接行动 - 信念一致性,用证据拒绝不可靠的强制消耗干预,并将策略效应与负载混淆分离。我们将外部信念在高噪声隐藏信息游戏中主要定位为可审计的认知基线,它也携带与决策相关的信号,将不透明的智能体行为转化为可重放的证据,以实现更安全、可控的迭代。

英文摘要

Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under strict, code-level information isolation, and we build an auditable framework that maintains an external belief state over hidden roles, logs belief updates and belief-action deviations as structured evidence, and supports a defensive offline improvement loop that reviews bad cases before any strategy change. Across 1,080 frozen games spanning belief-disabled, active-belief, kernel-ablation, camp-restricted, consumption-policy, and high-load arms, and including a seed-paired A0/A1 comparison, the active-belief condition is associated with substantially better good-side outcomes: in the 200-seed A0/A1 comparison the good-side win rate rises from 0.205 to 0.390 (paired McNemar $χ^2 = 16.4$, $p < 0.001$), with fewer irreversible witch-poison errors. We do not, however, attribute this shift to belief content. Direct action-belief consistency is low ($\approx 0.21$), and giving belief only to the werewolves helps the good side more than giving it only to the good side, which argues against a simple holder-benefit account; we therefore report the effect as an association and treat its mechanism as unresolved. The contribution is the audit framework itself: it makes the effect measurable, exposes low direct action-belief consistency, rejects an unreliable forced-consumption intervention with evidence, and separates strategy effects from load confounds. We accordingly position external belief in high-noise hidden-information games primarily as an auditable cognitive baseline that also carries decision-relevant signal, turning opaque agent behavior into replayable evidence for safer, controlled iteration.

Comments29 pages, 8 figures, 3 tables. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑