发表机构
School of Cyber Science and Engineering, Wuhan University; Shanghai Innovation Institute(武汉大学网络安全学院; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建多模态动态立场分类基准MMDS-Bench,评估12个多模态大语言模型,发现其在关系推理类多模态动态立场理解上仍存在不足。
AI 中文摘要
动态立场分类研究的是回复如何响应对其直接父消息,而非帖子与固定话题的关联方式。现有研究主要在纯文本场景中探讨该问题,而社交媒体互动越来越依赖图像、截图、表情包、反应图像及跨模态引用。我们推出MMDS-Bench,这是一个针对社交媒体父-回复互动中多模态动态立场分类的诊断基准。MMDS-Bench包含3482个多模态实例,采用七标签动态立场分类法标注,另有一个800实例的诊断子集,需对父消息理解、回复理解及立场关系推理进行结构化推理。我们还为每个实例标注了五个挑战因素,涵盖多模态融合、父消息框架、非字面表达、互动推理及标签边界模糊。我们评估了12个闭源和开源多模态大语言模型,并提出一种基于参考的LLM评判协议以评估推理质量。结果表明,当前多模态大语言模型(MLLMs)在多模态动态立场理解方面仍存在困难,尤其是在需要超越父消息与回复单独理解的关系推理场景中。
英文摘要
Dynamic stance classification models how a reply responds to its direct parent message, rather than how a post relates to a fixed topic. Existing work has mainly studied this problem in text-only settings, while social media interactions increasingly rely on images, screenshots, memes, reaction images, and cross-modal references. We introduce MMDS-Bench, a diagnostic benchmark for multimodal dynamic stance classification in social media parent-reply interactions. MMDS-Bench contains 3,482 multimodal instances annotated with a seven-label dynamic stance taxonomy, together with an 800-instance diagnostic subset that requires structured reasoning over parent understanding, reply understanding, and stance-relation inference. We further annotate each instance with five challenge factors covering multimodal fusion, parent framing, non-literal expression, interaction reasoning, and label-boundary ambiguity. We evaluate 12 closed-source and open-source multimodal large language models and propose a reference-grounded LLM-judge protocol for assessing reasoning quality. Results show that current MLLMs still struggle with multimodal dynamic stance understanding, especially in cases that require relational inference beyond separate parent and reply comprehension.
CommentsAccepted by EMNLP 2026