SpEmoC:一个平衡的说话者片段多模态情感基准
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
- Indian Institute of Technology Ropar(印度理工学院罗帕尔分校)
- Østfold University College(东福尔郡大学学院)
- Trinity College Dublin(都柏林圣三一大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究口语对话中人类情感理解,提出SpEmoC基准,整合预训练模型与人工验证标注七种情感模态,采用严格分割保持情感平衡分布,实验证明其能提升模型跨数据集评估时情感性能的稳定性。
AI中文摘要:
理解口语对话中的人类情感是情感计算中的关键挑战,在共情人工智能、人机交互和心理健康监测中有应用。现有数据集在规模、情感分布、模态对齐和数据划分策略上存在差异,影响跨数据集泛化和少数情感建模。我们引入SpEmoC,它包含来自3100部英语电影和电视剧的306,544个原始片段。从中精心挑选出30,000个高质量、类别平衡的片段,通过整合预训练模型和人工验证的混合管道标注七种情感的视觉、音频和文本模态。SpEmoC采用严格的电影和剧集级分割,保持七种情感的近乎平衡分布。实验表明平衡数据和精心分割能使模型在其他数据集上评估时情感性能更稳定,凸显了数据集设计对稳健可转移多模态情感识别的重要性。
英文摘要:
Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can influence reliable cross-dataset generalization and minority-emotion modeling. We introduce SpEmoC a Speaking segment Emotion for Conversations comprising 306,544 raw clips from 3,100 English language movies and TV series. From these, 30,000 high quality, class balanced clips are curated, featuring synchronized visual, audio, and textual modalities annotated for seven emotions through a hybrid pipeline that integrates pretrained models with human validation. SpEmoC uses strict movie- and series-level splits to prevent content overlap between split sets, allowing more reliable evaluation of model generalization. The dataset also maintains a near-balanced distribution across seven emotions, including minority classes such as Fear and Disgust, which supports more balanced learning across categories. Extensive experiments, including in-domain benchmarking, cross-dataset transfer, low-data training, class-imbalance analysis, and modality transfer show that balanced data and careful splitting lead to more stable performance across emotions when models are evaluated on other datasets. These results highlight the importance of dataset design for robust and transferable multimodal emotion recognition.