发表机构
College of AI, Tsinghua University; Department of Biostatistics and Bioinformatics, Duke University(清华大学人工智能学院; 杜克大学生物统计学与生物信息学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多方会议中多模态大语言模型的心理理论推理,引入MeetingToM基准,分层评估不同粒度社会行为推理,提供评估协议并分析模型,揭示其在整合线索、推断态度及区分共识方面的局限,为推进多模态模型的会议心理理论推理提供试验台。
AI 中文摘要
心理理论(ToM),即推断他人信念、意图和知识状态的能力,是社交互动的核心,但对当前多模态大语言模型(MLLMs)来说仍具挑战性,尤其是在多方会议中,线索分布于言语和行为之中。现有多模态ToM基准主要聚焦基于视频的公开、可外部验证信号的问答,对潜在社会状态和群体动态的覆盖有限。我们引入MeetingToM,一个用于自然主义多方会议中复杂社会行为推理的基准。MeetingToM针对特定会议现象,如“伪共识”。该基准分层组织,以评估不同社会粒度水平的ToM,包括主体级心理状态预测、二元级受话人理解和群体级共识推理。我们提供统一评估协议并对代表性MLLMs进行系统分析,揭示了在整合非语言线索、推断隐藏态度以及区分真共识与伪共识方面的持续局限性。我们的结果突出了关键挑战,并将MeetingToM确立为推进多模态模型中基于会议的ToM的试验台。
英文摘要
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics. We introduce MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings. MeetingToM targets meeting-specific phenomena such as \textbf{pseudo-consensus}, where apparent agreement masks private dissent under social pressure. The benchmark is hierarchically organized to evaluate ToM at increasing levels of social granularity, including (i) subject-level mental state prediction, (ii) dyadic-level addressee understanding, and (iii) group-level consensus reasoning. We provide a unified evaluation protocol and conduct systematic analyses of representative MLLMs, revealing persistent limitations in integrating non-verbal cues, inferring hidden attitudes, and distinguishing genuine consensus from pseudo-consensus. Our results highlight key challenges and establish MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.