Expresso-AI: 基于可解释视频的深度学习模型用于抑郁症诊断
Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis
- Media Lab MIT Cambridge, USA(媒体实验室 MIT 摄影实验室 美国剑桥)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出一种可解释的深度学习框架,通过微调动作识别预训练的DCNN,分析面部视频中的时空特征,生成显著性图以解释模型决策,同时提升抑郁症严重程度预测性能。
AI中文摘要:
鉴于抑郁症的广泛流行及其对个人和社会的重大影响,获得客观的早期诊断和干预措施至关重要。作为一个多学科课题,这些客观措施应具有可解释性,并可供医疗专业人员使用,以确保在心理健康护理领域进行有效的协作和治疗规划。尽管当前的自动抑郁症诊断方法在过去十年中有所改进,但仍存在关键差距,因为它们通常缺乏情感特异性和可解释性,限制了其实际应用和对心理健康护理的潜在影响。特别是,当使用深度模型时,视频中时间活动的可解释性尚未得到充分探索。在本研究中,我们提出了一种新颖的框架,用于分析深度神经网络在面部视频训练时的决策,特别关注自动抑郁症严重程度诊断。通过在AVEC抑郁症数据集的面部视频上微调在动作识别数据集上预训练的深度卷积神经网络(DCNN),我们的框架能够通过检查面部区域和时间表达语义来解释模型的显著性图。我们的方法为模型决策生成了视觉和定量解释,提供了对其推理过程的更深入洞察。除了这种可解释性,我们的基于视频的建模改进了先前用于视觉抑郁症诊断的单帧基准,从而提高了预测性能。总体而言,我们的工作展示了成功开发一个能够从面部模型决策中生成假设,同时提高抑郁症预测能力的框架。
英文摘要:
Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for early diagnosis and intervention. As a multidisciplinary topic, these objective measures should be interpretable and accessible to health care professionals, ensuring effective collaboration and treatment planning in the realm of mental health care. Even though current automated depression diagnosis approaches improved over the last decade, a critical gap exists as they often lack affect-specificity and interpretability, limiting their practical application and potential impact on mental health care. In particular, interpretability from temporal activities from videos when deep models are used is not fully explored. In this study, we present a novel framework for analyzing Deep Neural Networks' decisions when trained on facial videos, specifically focusing on automatic depression severity diagnosis. By fine-tuning Deep Convolutional Neural Networks (DCNN) pre-trained on Action Recognition datasets on depression severity facial videos from AVEC depression dataset, our framework is able to interpret the model's saliency maps by examining face regions and temporal expression semantics. Our approach generates both visual and quantitative explanations for the model's decisions, providing greater insight into its reasoning. In addition to this interpretability, our video-based modeling has improved upon previous single-face benchmarks for visual depression diagnosis, resulting in enhanced predictive performance. Overall, our work demonstrates the successful development of a framework capable of generating hypotheses from a facial model's decisions while simultaneously improving depression's predictive capabilities.