arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VIP-MINGLE:用于群体语言交流中视频会议和面对面多模态交互的语料库

VIP-MINGLE: A Corpus for Videoconference and In-Person Multimodal Interaction in Group Language Engagement

Andrew Chang, Abhinay K Bodi, Wenxin Deng, Junrui Huang, Venu G Kadamba, Sumanth B H Karanam, Dhiwahar A Kennady, David Poeppel, Dustin Freeman

arXiv 2607.13614首次发表:更新:

AI 中文总结

介绍VIP-MINGLE多模态数据集,含59小时录音等,有两种环境下配对会话,涵盖多种数据。分析发现不同环境多模态行为分布有变化,该数据集是跨环境群体对话模型开发的关键资源。

AI 中文摘要

群体对话是社会互动的一种基本但复杂的形式,对人类认知和电信技术至关重要。尽管理解和促进这些互动一直是长期目标,但由于缺乏连接两者的数据集,研究结果往往局限于特定的面对面或视频会议环境。我们引入了VIP-MINGLE,这是一个多模态数据集,包含59小时的录音(32个小组,105名参与者),具有两种环境下的配对主体会话。该数据集包括原始音频/视频、心理测量数据、处理后的多模态特征(如分帧语音、面部表情、转录)和时间分辨的人类注释。我们的分析揭示了不同环境下多种模态之间显著的行为分布变化,凸显了跨环境语料库的必要性。VIP-MINGLE是开发跨环境群体对话强大模型的关键资源。

英文摘要

Group conversations are a fundamental yet complex form of social interaction central to human cognition and telecommunication technology. While understanding and facilitating these interactions has been a long-standing goal, findings are often isolated within specific in-person or videoconferencing settings due to a scarcity of datasets that bridge the two. We introduce VIP-MINGLE, a multimodal dataset comprising 59 hours of recordings (32 groups, 105 participants), featuring paired within-subject sessions in both settings. The dataset includes raw audio/video, psychometric data, processed multimodal features (e.g., diarized speech, facial expressions, transcriptions), and time-resolved human annotations. Our analysis reveals significant behavioral distribution shifts across multiple modalities between settings, reinforcing the need for a cross-setting corpus. VIP-MINGLE serves as a critical resource for developing robust models of group conversations across settings.

CommentsInterspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑