无行为的信念:测量视觉-语言模型中心智理论到协同社会行动的转化
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
该研究提出基准MOSAIC评估视觉-语言模型的心智理论到协同社会行动的转化,发现多数模型存在瓶颈,而带显式ToM模块的PCM-LLM表现优异,证实显式信念-行动耦合的作用。
中文摘要 AI 辅助
有效的社会互动要求智能体同时将心理状态推理转化为跨言语和非言语渠道的协同行为信号。然而现有基准单独评估心智理论(ToM)推理和具身行为,未测量社会推理与社会行动之间的差距。我们引入MOSAIC(Multimodal Orchestration of Social Action, Inference, and Communication,即社会行动、推理与沟通的多模态编排),这是一个受控基准,其中两个具身智能体在合作与竞争场景中互动,这些场景需要在系统变化的ToM约束下整合言语陈述、空间轨迹、注视方向和面部表情。我们对13个模型(包括11个视觉-语言模型(VLM))进行评估,每个模型完成200次试验,结果发现VLM无法在ToM顺序约束下产生与预期结果一致的行为,且施加显式ToM顺序约束不会产生与指定推理水平一致的可靠行为变化。信号级分析揭示了两个连续瓶颈:大多数模型无法产生方向连贯的非言语信号,即使存在信号,VLM智能体也无法解释其他智能体的行为并做出反应。作为具有显式ToM模块的结构化架构参考点的PCM-LLM,在所有条件下均取得成功,表明显式信念-行动耦合是这类任务的充分要素。
英文摘要
Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.
发表机构
- CIAMS, Université Paris-Saclay(巴黎萨克雷大学 CIAMS)
- CQSB, Sorbonne Université(索邦大学 CQSB)
机构由 AI 辅助整理,请以论文原文为准。