AI 中文总结
介绍VIP-MINGLE多模态数据集,含59小时录音等,有两种环境下配对会话,涵盖多种数据。分析发现不同环境多模态行为分布有变化,该数据集是跨环境群体对话模型开发的关键资源。
AI 中文摘要
群体对话是社会互动的一种基本但复杂的形式,对人类认知和电信技术至关重要。尽管理解和促进这些互动一直是长期目标,但由于缺乏连接两者的数据集,研究结果往往局限于特定的面对面或视频会议环境。我们引入了VIP-MINGLE,这是一个多模态数据集,包含59小时的录音(32个小组,105名参与者),具有两种环境下的配对主体会话。该数据集包括原始音频/视频、心理测量数据、处理后的多模态特征(如分帧语音、面部表情、转录)和时间分辨的人类注释。我们的分析揭示了不同环境下多种模态之间显著的行为分布变化,凸显了跨环境语料库的必要性。VIP-MINGLE是开发跨环境群体对话强大模型的关键资源。
英文摘要
Group conversations are a fundamental yet complex form of social interaction central to human cognition and telecommunication technology. While understanding and facilitating these interactions has been a long-standing goal, findings are often isolated within specific in-person or videoconferencing settings due to a scarcity of datasets that bridge the two. We introduce VIP-MINGLE, a multimodal dataset comprising 59 hours of recordings (32 groups, 105 participants), featuring paired within-subject sessions in both settings. The dataset includes raw audio/video, psychometric data, processed multimodal features (e.g., diarized speech, facial expressions, transcriptions), and time-resolved human annotations. Our analysis reveals significant behavioral distribution shifts across multiple modalities between settings, reinforcing the need for a cross-setting corpus. VIP-MINGLE serves as a critical resource for developing robust models of group conversations across settings.
CommentsInterspeech 2026