多模态时间建模用于多方对话中的连续群体情绪识别
Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues
浏览论文内容
中文总结 AI 辅助
本研究提出一秒分辨率的多模态时间框架,利用滑动窗口整合音视频,在TEIDAN数据集上实现连续群体情绪识别,并引入混合状态,实验表明时间Transformer优于基线且揭示情绪分歧为主要挑战。
中文摘要 AI 辅助
为了在多方位对话场景中实现对话智能体的自然行为,理解群体情绪(如效价和唤醒度)的整体状态至关重要。以往大多数工作是在话语层面或使用粗粒度时间窗口来处理这一任务,这不足以捕捉情绪动态。在本研究中,我们将群体情绪的连续识别设定为一秒分辨率。此外,我们还引入了混合状态,该状态捕捉群体中参与者之间的情绪分歧。我们基于TEIDAN语料库构建了一个具有帧级软标签的数据集,并提出了一种多模态时间框架,利用滑动窗口上下文整合音频和视频信息。实验结果表明,时间Transformer优于简单基线,并且与基于LLM的模型相比,与地面真值标签表现出更强的时间一致性。上下文长度的影响有限,而视听输入在连续标签指标上优于任一单模态输入。此外,我们的分析显示,在混合值较高的时间区间内,群体情绪识别误差较大,这揭示了情绪分歧是群体情绪识别的关键挑战。
英文摘要
To realize natural behavior in dialogue agents in multi-party dialogue scenarios, it is important to understand group emotion such as valence and arousal as a whole. Most prior work addressed this task at the utterance level or using a coarse-grained time window, which is not sufficient to capture emotional dynamics. In this study, we formulate continuous recognition of the Group Emotion at a one-second resolution. Moreover, we also introduce the Mixed state, which captures the emotional divergence among participants in the group. We constructed a dataset with frame-level soft labels based on the TEIDAN corpus and propose a multimodal temporal framework that integrates audio and video information using a sliding-window context. Experimental results demonstrate that the temporal Transformer outperforms simple baselines and shows stronger temporal agreement with the ground-truth labels than the LLM-based model. The effect of context length is limited, whereas audio-visual input outperforms either unimodal input on the continuous-label metrics. Additionally, our analysis shows larger Group Emotion recognition errors in intervals with high Mixed values, exposing emotional divergence as a key challenge for group emotion recognition.
发表机构
- Kyoto University(京都大学)
机构由 AI 辅助整理,请以论文原文为准。