arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30716cs.CLcs.CV

SocialReasonBench:基于反事实叙事视频的社会推理视频问答基准

SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos

Zheyu Huang, Zijing Shi, Haozhe Luo, Huadong Tang, Mingyu Liu, Meng Fang, Ling Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员推出SocialReasonBench基准,评估LMMs的社会推理能力,发现其在反事实和因果推理上表现不佳,凸显了模型在潜在社会状态推理上的不足。

中文摘要 AI 辅助

大型多模态模型(LMMs)的最新进展大幅提升了视频理解能力,但其对以人为中心的社会情境的推理能力仍然有限。现有基准通常依赖于仅包含单一观测轨迹的视频,难以判断模型是真正理解社会动态,还是仅仅利用重复出现的叙事模式。我们推出SocialReasonBench,这是一个视频多项选择问答基准,用于评估交互式叙事场景中基于社会基础的推理能力。该基准基于《底特律:变人》(Detroit: Become Human)的游戏玩法视频构建,利用分支故事情节,其中玩家决策会导致不同的社会结果,这些结果可与游戏自身的脚本、流程图和记录的分支进行核对。我们开发了一个多智能体策划流程,该流程可定位具有社会意义的片段,基于游戏状态信号确定答案标签,并生成具有诊断性干扰项的理论指导问题。SocialReasonBench涵盖七个推理维度,包括意图识别、情感共情、道德困境、反事实推理和因果前件等。对当代LMMs的实验表明,模型在基础社会理解上表现尚可,但在反事实和因果推理方面存在困难。进一步的消融研究和诊断错误分析显示,模型往往依赖不完整的模态线索,并陷入视觉捷径等推理陷阱,凸显了可观测事件识别与对潜在社会状态进行更深层次推理之间的差距。

英文摘要

Recent advances in Large Multimodal Models (LMMs) have greatly improved video understanding, yet their ability to reason about human-centered social situations remains limited. Existing benchmarks typically rely on videos with a single observed trajectory, making it difficult to determine whether models truly understand social dynamics or merely exploit recurring narrative patterns. We introduce SocialReasonBench, a video multiple-choice QA benchmark for evaluating socially grounded reasoning in scenarios derived from interactive narratives. Built from gameplay videos of Detroit: Become Human, the benchmark leverages branching storylines where player decisions lead to alternative social outcomes that can be checked against the game's own script, flowchart, and recorded branches. We develop a multi-agent curation pipeline that localizes socially meaningful clips, grounds answer labels in game-state signals, and generates theory-guided questions with diagnostic distractors. SocialReasonBench covers seven reasoning dimensions, including intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent. Experiments on contemporary LMMs show that models perform reasonably well on basic social understanding but struggle with counterfactual and causal reasoning. Further ablation and diagnostic error analyses reveal that models often depend on incomplete modality cues and fall into reasoning traps such as visual shortcuts, highlighting a gap between observable event recognition and deeper reasoning over latent social states.

发表机构

  • AAII, University of Technology Sydney(悉尼科技大学AAII(先进人工智能研究所))
  • Northeastern University(东北大学)
  • Fudan University(复旦大学)
  • University of Liverpool(利物浦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑