arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ASD-FEAT:用于早期自闭症谱系障碍风险预测的多模态婴儿视频数据集

ASD-FEAT: A Multi-Modal Infant Video-Derived Dataset for Early ASD Risk Prediction

Sidrah Liaqat, Halil Helvaci, Sen-Ching Cheung, Chongruo Wu, Dongjie Chen, Chen Nee Chuah, Sally Ozonoff

arXiv 2610.04051首次发表:更新:

发表机构

University of Kentucky(肯塔基大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ASD-FEAT多模态婴儿视频数据集,结合纵向覆盖、专家标注与隐私特征,并基于计算机视觉流程实现76.2%准确率、0.82 AUROC的ASD风险自动预测,引入伙伴对比进一步提升性能。

AI 中文摘要

自闭症谱系障碍(ASD)的准确早期筛查是及时干预的前提,而及时干预对于改善认知和行为结果至关重要。我们提出了ASD-FEAT(ASD - 特征提取与追踪),这是一个从婴儿与成人互动视频记录中衍生的多模态数据集。ASD-FEAT的关键贡献在于将以下特性相结合:从婴儿期到36个月的纵向覆盖、重复的互动会话、经临床验证的发展结果、专家帧级行为标注,以及注重隐私的多模态特征表示。据我们所知,现有的ASD行为数据集无法在同等规模上同时提供这些特性。为展示ASD-FEAT的实用性,我们利用它来评估一个基于计算机视觉的端到端流程,该流程依赖机器学习技术自动识别ASD风险。ASD-FEAT整合了专家定义的特征和深度学习特征,包括面部和眼部关键点、面部动作单元、注视方向、头部位置、梅尔频谱图音频表示和光流,以识别社会互动的行为标记。我们的自动化流程实现了76.2%的ASD分类准确率和0.82的受试者工作特征曲线下面积(AUROC),而基于人工标注行为训练的分类器则达到了81.3%的准确率和0.88的AUROC。我们进一步引入了访问内伙伴对比:一种每次访问的信号,对比检查者导向和父母导向的社会行为,当将其添加到分类器中时,将全自动的“看脸+微笑”和“看脸+发声”配置的马修斯相关系数分别提升至0.49和0.45,超过了人工编码的单伙伴基线0.42。

英文摘要

Accurate early screening for Autism Spectrum Disorder (ASD) is a precursor to timely intervention, which is critical for improving cognitive and behavioral outcomes. We present ASD-FEAT (ASD - Feature Extraction And Tracking), a multimodal dataset derived from video recordings of infant-adult interaction sessions. The key contribution of ASD-FEAT is the combination of longitudinal coverage from infancy through 36 months, repeated interaction sessions, clinically validated developmental outcomes, expert frame-level behavioral annotations, and privacy-conscious multimodal feature representations. To the best of our knowledge, existing ASD behavioral datasets do not jointly provide these characteristics at comparable scale. To demonstrate the utility of ASD-FEAT, we use it to evaluate a computer-vision-based end-to-end pipeline relying on machine learning techniques to automatically identify ASD risk. ASD-FEAT integrates both expert-defined and deep-learned features, including face and eye landmarks, facial action units, gaze direction, head position, mel-spectrogram audio representations, and optical flow, to identify behavioral markers of social interaction. Our automated pipeline achieves an ASD classification accuracy of 76.2% and an Area Under the Receiver Operating Characteristic (AUROC) of 0.82, compared to classifiers trained on manually labeled behaviors, which yielded 81.3% accuracy and an AUROC of 0.88. We further introduce a within-visit partner contrast: a per-visit signal contrasting examiner-directed and parent-directed social behavior which, when added to the classifier, lifts the fully automated Look Face + Smile and Look Face + Vocal configurations to Matthews correlations of 0.49 and 0.45 respectively, exceeding the human-coded single-partner baseline of 0.42.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑