Causal-EVC:通过时空基础与反事实干预打破情感虚假因果
Causal-EVC: Breaking Emotional Spurious Causality via Spatiotemporal Grounding and Counterfactual Intervention
浏览论文内容
中文总结 AI 辅助
针对情感视频字幕中的虚假因果与共现偏差问题,提出Causal-EVC框架,通过时空定位模块和反事实对比目标,在构建的EVC-CauseGround基准上提升因果忠实性,实现可解释的多模态情感理解。
中文摘要 AI 辅助
情感视频字幕生成旨在生成事实准确且富有情感共鸣的描述。尽管近期方法已认识到视觉原因在引导情感感知和字幕生成中的重要性,但它们从根本上依赖于简单的注意力匹配,这不可避免地会遭受共现偏差中的因果冗余和虚假相关(例如,在阳光明媚的海滩上将“悲伤”误分类为“喜悦”),导致从混淆背景中产生严重的捷径学习。此外,现有评估未能验证模型是否真正掌握了因果推理,或仅仅利用了背景混淆因素。为解决这些局限,我们首先构建了EVC-CauseGround,一个具有密集时空因果标注的综合基准。关键的是,它引入了一个精心挑选的因果忠实子集,以明确量化真实的情感-原因归因。其次,我们提出了Causal-EVC,一个情感基础字幕生成框架,该框架引入了一个运动引导的因果时空定位模块,以精确地将因果触发因素从背景混淆因素中解耦。此外,我们引入了一个可解释的稀疏情感路由模块。通过合成反事实表示并制定一个新的反事实对比目标,我们强制模型将其情感预测严格锚定在真实的因果触发因素上,而非混淆背景。大量实验表明,Causal-EVC不仅在语义指标上取得了最佳性能,而且在因果忠实子集上也表现出显著优势,这证明我们的模型能够从真实的视觉原因中挖掘情感线索,并减轻共现偏差,以实现可解释的多模态情感理解。
英文摘要
Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple attention matching, which inevitably suffers from {causal redundancy and spurious correlations} in co-occurrence bias (e.g., misclassifying ``sadness'' as ``joy'' on a sunny beach), leading to severe shortcut learning from confusing backgrounds. Furthermore, existing evaluations fail to verify whether models have genuinely mastered causal reasoning or merely exploited background confounders. To address these limitations, we first construct {EVC-CauseGround}, a comprehensive benchmark with dense spatio-temporal causal annotations. Crucially, it introduces a carefully selected {Causal-Faithfulness Subset} to explicitly quantify genuine emotion-cause attribution. Second, we propose {Causal-EVC}, an emotion-grounding captioning framework, which introduces a Motion-guided Causal Spatiotemporal Localization module to precisely decouple causal triggers from background confounders. Besides, we introduce an Interpretable Sparse Emotion Routing module. By synthesizing counterfactual representations and formulating a novel counterfactual contrastive objective, we enforce the model to anchor its emotion predictions strictly on authentic causal triggers instead of confusing background. Extensive experiments show that Causal-EVC not only achieves the best performance on semantic metrics but also exhibits significant advantages in the causal-faithfulness subset, which demonstrates that our model could mine emotional cues from genuine visual causes and mitigate co-occurrence bias for interpretable multimodal emotion understanding.
发表机构
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。