arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24191cs.CLcs.AI

StanceFlip:用于多模态对话立场翻转预测的综合多维基准测试

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

Heyan Chai, Xin Li, Wenjie Wang, Jianyang Qin, Chaoyang Li, Lu Wang, Hao Chen, Qing Liao

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现有对话立场检测基准测试的局限,提出StanceFlip基准测试及ConStaFF框架,含多模态立场六元组提取和动态立场翻转归因两个新子任务,基于大语言模型实现端到端立场推理,实验显示该方法性能领先。

中文摘要 AI 辅助

对话立场检测已从静态文本分析转向动态多模态建模。然而,现有基准测试存在三个关键限制:无法捕捉信念的动态演变,尤其是在立场反转期间;难以将情感状态与逻辑推理区分开;忽视多模态线索在解决语用歧义(如讽刺)中的关键作用。为解决这些限制,我们提出了StanceFlip,这是一个用于多轮对话中跨五种模态和多场景的多模态对话立场翻转预测的基准测试,包括两个新子任务:多模态立场六元组提取,提取持有者、目标、情感、情绪、立场和基本原理作为对话的静态状态快照以捕捉细粒度认知结构;动态立场翻转归因,跟踪对话中的立场反转并识别其潜在触发因素。除数据集外,我们还提出了一个名为ConStaFF的专用框架用于多模态对话立场翻转预测。基于大语言模型构建,ConStaFF执行端到端立场推理,集成了立场思考推理框架和自我反思验证机制以进行结构化立场建模和可靠的翻转归因。具体而言,立场思考将推理过程分解为专门的认知角色以制定目标命题、解决跨模态冲突并推断历史立场轨迹。大量实验表明,我们的方法在六元组提取和翻转触发归因方面均取得了领先性能,大幅优于强大得多模态大语言模型基线。

英文摘要

Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed for multimodal conversational stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel subtasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Forecasting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mechanism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, resolve cross-modal conflicts, and infer historical stance trajectories. Extensive experiments show that our approach achieves state-of-the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins.

发表机构

  • College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件学院)
  • Harbin Institute of Technology Shenzhen(哈尔滨工业大学(深圳))
  • City University of Macau(澳门城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑