EmotionDialogCN:面向普通话情感对话的自发多模态数据集
EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue
AI总结:
该研究推出EmotionDialogCN数据集,其规模大、情感分布贴合真实情况,采用新型采集框架减少设备干扰,可支撑多模态情感相关研究。
AI中文摘要:
面对面视听交互是人类交流的核心,能传递丰富的情感与社交线索。但现有多模态对话数据集存在情感标注不足、情感多样性差、规模小的局限。我们推出EmotionDialogCN,这是一个旨在捕捉真实面对面交流的大规模视听情感数据集,包含119名专业演员在20种日常场景中完成的21880段对话会话,涵盖18种情感类别,录音时长超400小时,是同类中规模最大、最全面的数据集。一种新型数据采集框架最大程度减少设备干扰,使情感表达自然且细腻。EmotionDialogCN的情感分布与真实人类情感统计的偏差为0.64,而此前数据集的该值为5.65,且主体构图一致(帧占比52%-59%)。这些特性共同实现了声学、词汇、视觉模态下单模态与多模态的稳定表现,融合结果进一步凸显了强大的多模态对齐与跨模态互补性。
英文摘要:
Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by inadequate emotion annotations, poor emotional diversity, and small scale. We introduce EmotionDialogCN, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication. It contains 21,880 dialogue sessions performed by 119 professional actors across 20 everyday scenarios, covering 18 emotion categories with over 400 hours of recordings, the largest and most comprehensive dataset of its kind. A novel data collection framework minimizes equipment interference, enabling natural and nuanced emotional expressions. EmotionDialogCN achieves an emotion distribution deviation of 0.64 from real human emotion statistics (versus 5.65 for prior datasets) and consistent subject framing (52-59% frame occupancy). Together, these properties translate into stable unimodal and multimodal performance across acoustic, lexical, and visual modalities, with fusion results further underscoring strong multimodal alignment and cross-modal complementarity.