arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24146cs.LG

心智还是言语?多智能体社会模拟中的心智理论审计

Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation

Cong Li, Cheng Chen, Thomas Fung, Alex Rossi, Yi Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过构建160个二元组谈判模拟,使用2880个反事实探针审计语言模型智能体的心智理论,发现其社交流畅但经济表现差,依赖自我中心投射而非真正推断伙伴心智。

中文摘要 AI 辅助

语言模型智能体越来越多地被用于模拟社会互动,所产生的对话记录读起来仿佛智能体彼此理解。我们探究这种表象是建立在对伙伴心智的模型之上,还是仅基于伙伴所说内容的表面记录。我们构建了一个社会模拟,其中两个问题都有精确答案:40场多议题谈判,其隐藏的偏好权重和完整的帕累托前沿由构造已知。两个模型家族在160个二元组中进行谈判,每份对话记录在测量前被冻结,随后2880个反事实探针在保持证据字节完全相同的同时,一次只变动一个因素:读者自身的利益、伙伴的语气、身份标签以及递归顺序。这些智能体在社交上流畅但在经济上表现不佳。它们在96.2%的二元组中达成协议,且无协议失败,但只有0.7%的交易落在帕累托前沿上,它们留下了20.5%的可用联合价值未获取,并且在76.6%的交易中错过了双方利益完全一致的那个议题;在前沿和该一致议题上,从双方都会接受的集合中随机抽取的包裹表现相同。探针定位了失败原因。仅交换读者自身的收益表,而伙伴的言语和提议保持不变,会使推断出的首要优先级变动15.0个百分点,这是自我中心投射而非推断,而语气改写使其变动5.3个百分点,身份标签变动0.0。最值得注意的是,智能体在72.5%的时间里预测其伙伴对它的信念,而该伙伴的信念本身只有51.2%的时间是正确的:智能体跟踪对话的能力远好于跟踪对话背后的心智。

英文摘要

Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both questions have exact answers: 40 multi-issue negotiations whose hidden preference weights and whose full Pareto frontier are known by construction. Two model families negotiate across 160 dyads, every transcript is frozen before any measurement, and 2880 counterfactual probes then hold the evidence byte identical while moving one factor at a time: the reader's own stake, the partner's tone, an identity label, and the order of recursion. The agents are socially fluent and economically poor. They reach agreement in 96.2% of dyads with 0 protocol failures, yet only 0.7% of deals land on the Pareto frontier, they leave 20.5% of the available joint value unclaimed, and they miss the one issue on which their interests are perfectly aligned in 76.6% of deals; on the frontier and on that aligned issue, a package drawn at random from the set both sides would accept does as well. The probes locate the failure. Swapping only the reader's own payoff sheet, while the partner's words and offers stay identical, moves the inferred top priority by 15.0 percentage points, which is egocentric projection rather than inference, while a tone rewrite moves it by 5.3 percentage points and an identity label by 0.0. Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.

发表机构

  • DailyAdvance

机构由 AI 辅助整理,请以论文原文为准。

↑