AI 中文总结
针对多模态语音分析领域因现有数据集存在局限而发展受限的问题,引入CARE v1.0多模态英语数据集,涵盖多种医疗状况,提供丰富描述符和元数据,支持多种应用,推动该领域研究发展。
AI 中文摘要
自动分析多模态语音在通过计算检测和监测多种神经、精神和呼吸疾病方面显示出强大潜力。然而,该领域进展受现有公开数据集限制,其规模小、专注单一疾病且主要聚焦语音。此外,关键混杂变量记录不足影响计算分析可靠性和可解释性。为解决这些问题,我们引入CARE v1.0,这是一个精心策划的多模态英语数据集,包含从612名个体收集的约144小时短视频访谈,涵盖12种医疗状况及一个对照队列。每个视频都提供了一套全面的临床相关多模态描述符以及结构化元数据。该语料库的广度和异质性支持广泛应用,包括自动疾病和症状检测、情绪激动情境下言语和非言语行为的多模态建模以及疾病轨迹和应对过程研究。
英文摘要
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are often small in scale, focused on a single condition or disease, and primarily speech focused. Moreover, if key confounding variables such as education, medication use, comorbidities, or mood state are insufficiently documented, the reliability and interpretability of computational analyses are further compromised. To address these limitations, we introduce CARE v1.0, a curated multimodal English dataset of approximately 144 hours of short video interviews collected from 612 individuals across 12 medical conditions plus a control cohort. For each video, a comprehensive set of clinically relevant multimodal descriptors is provided, alongside structured metadata covering factors such as medication, life impacts, and expressed emotions. The corpus's breadth and heterogeneity support a wide range of applications, including automatic disease and symptom detection, multimodal modelling of speech and non-verbal behaviour under emotionally charged contexts, and studies of disease trajectories and coping processes.
CommentsUnder review in npj Scientific Data