arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MafiaScope:社交推理游戏中对大语言模型智能体的非侵入性、时间分辨信念探测

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

Ilia Karpov

arXiv 2607.10645首次发表:更新:

发表机构

HSE University(俄罗斯高等经济研究大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究社交推理游戏中LLM智能体信念探测问题,提出MafiaScope测试平台,通过结构化问题让智能体私下作答并自动评分,可视化呈现信念轨迹,以DeepSeek为例进行研究,得出相关校准误差等结果并开源相关内容。

AI 中文摘要

大语言模型智能体的公开行为很难揭示其社交推理:正确投票的智能体可能是在猜测,善于说谎的智能体不会留下其真实信念的痕迹。我们提出了MafiaScope,一个将社交推理游戏《黑手党》转变为机器心理理论测量工具的开放测试平台。每次公开发言后,每个智能体私下回答一组可配置的结构化探测问题;答案不会重新进入游戏,并根据引擎所知的地面真值自动评分。交互式可视化工具呈现信念轨迹:模拟模式展示游戏在一个智能体视角下的情况,面板绘制与时间线对齐的准确性和校准情况,反事实回放可分叉任何记录步骤。在一个包含13815个解析探测答案的32场游戏的DeepSeek案例研究中,陈述的置信度校准不佳,预期校准误差为0.17,智能体过度预测被怀疑的次数为1.5倍,并且一个30分叉的回放实验从头到尾运行了反事实回放工作流程。引擎、查看器和200多个跨模型游戏的语料库在开放许可下发布;现场演示:此https URL 屏幕录像:此https URL。

英文摘要

An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testbed that turns the social deduction game Mafia into a measurement instrument for machine Theory of Mind. It distinguishes whether an agent lost because it misread the game or because it failed to act on a correct assessment, a distinction that is invisible from outcomes and dialogue transcripts alone. After every public utterance, each agent privately answers structured probe questions whose responses never re-enter the game and are scored against the ground truth known to the engine. An interactive visualizer replays games from the perspective of an individual agent's beliefs, displays timeline-aligned accuracy and calibration, and supports counterfactual replay from any recorded step. In a case study across two model families comprising tens of thousands of parsed probe responses, we find that stated confidence is poorly calibrated, agents overestimate how often they are suspected by a factor of 1.5, and single-vote counterfactual replays rarely change game outcomes: outcome flips occur primarily when the agent had already formed a correct belief state, whereas decisions made under an incorrect model of the world remain largely unchanged under resampling. The engine, visualizer, recorded games, and counterfactual replay corpus are released under an open-source licence. Code: https://github.com/karpovilia/mafiascope. Live demo: https://karpovilia.github.io/mafiascope/. Screencast: https://vimeo.com/1208920221.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑