发表机构
The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MultivationBench基准,基于心理学框架评估多模态序列动机推理,发现现有模型难以在序列语境中保持一致的动机推理,暴露了静态识别与动态推理的差距。
AI 中文摘要
多模态大语言模型因具备社会智能潜力引发广泛关注,但其序列动机推理能力仍未得到充分研究。现有评估主要针对静态文本或孤立视觉快照,无法反映现实世界行为驱动的累积性。为解决这一缺口,本文提出MultivationBench——一个用于严格评估故事驱动视觉叙事中多模态动机推理的基准。该基准基于马斯洛需求层次和赖斯基本欲望等成熟心理学框架构建,要求模型整合累积的多模态上下文以推断不断演变的动机。结果表明,MultivationBench构成重大挑战:所有测试模型均难以在序列语境中保持一致的动机推理,凸显了静态识别能力与类人社会理解所需动态推理之间的关键差距。
英文摘要
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.
Comments35 pages, 6 figures. Accepted to Findings of EMNLP 2026. Code and data: https://github.com/HKUST-KnowComp/MultivationBench