MoCA:隐式社会语境分析
MoCA: Implicit Social Context Analysis
- National University of Singapore(新加坡国立大学)
- University of Oxford(牛津大学)
- Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出隐式社会语境分析(MoCA)新任务,构建含3108个实例的基准数据集,提出CoDAR框架提升模型性能,但模型与人类推理仍存差距,凸显隐式社会理解难度。
AI中文摘要:
人类的社会交流,如情感与意图,常以高度隐式的方式传递,其潜在含义通过间接的、基于社会与文化的信号而非明确表述来表达。这类隐式社会语境在现实互动中普遍存在,但目前仍缺乏用于研究它们的正式且系统的框架。本文中,我们提出隐式社会语境分析(MoCA)这一新任务,该任务沿情感、意图和立场三个关键维度对隐式社会场景进行系统建模。我们构建了一个包含3108个多模态实例的高质量基准数据集,这些实例收集自现实来源,并带有细粒度的认知标注,标注内容包括谁向谁表达了什么,以及如何和为何进行表达。利用MoCA数据集,我们发现最先进的多模态大语言模型在该任务上表现显著不佳,因为它们依赖明确线索,且对潜在社会语境进行推理的能力有限。为应对这一挑战,我们提出冲突驱动溯因推理(CoDAR)这一新框架,该框架将观察到的表达与预期真实行为之间的差异建模为认知冲突,从而实现对隐藏心理状态的推理。大量实验表明,CoDAR大幅提升了模型性能,但与人类推理仍存在较大差距,凸显了隐式社会理解的根本难度。
英文摘要:
Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through indirect, socially and culturally grounded signals rather than explicit statements. Such implicit social contexts are pervasive in real-world interactions, yet there remains a lack of a formal and systematic framework for studying them. In this paper, we introduce Implicit Social Context Analysis (MoCA), a novel task that systematically models implicit social scenarios along three key dimensions: affection, intent, and stance. We construct a high-quality benchmark containing 3,108 multimodal instances collected from real-world sources, with fine-grained cognitive annotations revealing who expresses what toward whom, as well as how and why it is conveyed. Using the MoCA dataset, we show that state-of-the-art multimodal large language models struggle significantly with this task because of their reliance on explicit cues and limited ability to reason over latent social contexts. To address this challenge, we propose Conflict-Driven Abductive Reasoning (CoDAR), a novel framework that models the discrepancy between observed expressions and expected truthful behavior as cognitive conflict, thereby enabling the inference of hidden mental states. Extensive experiments demonstrate that CoDAR substantially improves model performance. Nevertheless, a large gap from human reasoning remains, highlighting the fundamental difficulty of implicit social understanding.