Avalon-ToM-Bench:通过非对称游戏机制评估细粒度心理理论
Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics
- National Taiwan University(台湾大学)
- CyCraft AI Lab(CyCraft人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Avalon-ToM-Bench基准,基于阿瓦隆游戏机制评估大语言模型的细粒度心理理论,发现模型ToM能力不足源于推理策略而非知识或表征,推理训练比测试时思维链增益更显著。
AI中文摘要:
心理理论(Theory of Mind,ToM)是智能体交互的关键,但现有评估要么依赖过度简化心理状态推理的静态场景,要么提供有限诊断视角的交互设置。本文提出Avalon-ToM-Bench,这一细粒度基准通过《抵抗:阿瓦隆》(The Resistance: Avalon)的非对称信息机制将ToM落地实现。该基准不评估端到端游戏玩法,而是将ToM分解为2×2分类——认知与动机推理的交叉,结合推理与行动维度,采用人工构建的视角约束查询。对28个大语言模型(LLM)的基准测试揭示了三点洞见:1. 是推理而非知识:模型展现出较强的游戏规则理解能力,但ToM能力明显较弱,问题出在社会推理而非领域知识缺失;2. 是表达而非表征:通过线性探测和激活调控的机制分析显示,模型的隐藏状态中常包含正确的心理状态推理表征,但在生成时无法有效表达——线性探测的准确率达77%-82%,而模型自身的思维链仅为62%-70%;3. 是策略而非深思熟虑:专门的推理训练带来显著提升,而测试时思维链仅提供微小增益(平均分别提升11.0分与1.1分),表明稳健的ToM依赖习得的推理策略而非推理时的额外深思熟虑。
英文摘要:
Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a 2$\times$2 taxonomy -- epistemic versus motivational reasoning crossed with inference versus action -- using human-crafted, perspective-constrained queries. Benchmarking 28 LLMs reveals three insights: 1) Reasoning, not knowledge. Models show strong game-rule comprehension but markedly weaker ToM abilities, isolating failures to social reasoning rather than missing domain knowledge. 2) Expression, not representation. Mechanistic analyses via linear probing and activation steering show that models frequently represent correct mental-state inferences in their hidden states but fail to express them during generation -- linear probes recover 77-82% accuracy versus 62-70% from the models' own chain-of-thought. 3) Policy, not deliberation. Dedicated reasoning training yields substantial improvements whereas test-time chain-of-thought provides only marginal gains (+11.0 versus +1.1 points on average), suggesting that robust ToM depends on a learned reasoning policy rather than increased inference-time deliberation.