arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22794cs.LGcs.AIcs.CLcs.SD

用于抑郁症检测的多模态域泛化:基于注意力的双向长短期记忆网络与域对抗训练

Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training

Ali Tabaraei, Federico Simonetta, Stavros Ntalampiras

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对深度学习抑郁症检测泛化受限问题,提出含域泛化的多模态检测框架,集成BiLSTM与注意力机制,用梯度反转层增强泛化,实验表明该方法提升了准确率和F1分数,超越现有基准,消融研究突出各因素贡献。

中文摘要 AI 辅助

利用深度学习进行自动抑郁症检测虽有前景,但因说话者间差异导致的域转移,泛化能力受限。本文提出首个独立于患者的多模态抑郁症检测框架,融合域泛化(DG),联合利用声学和文本模态。模型将双向长短期记忆网络(BiLSTM)与模态内和跨模态注意力机制集成,通过段级融合决策。受神经网络域对抗训练(DANN)启发应用梯度反转层,增强泛化能力,减少患者特定偏差。在Androids-Corpus数据集上用5折交叉验证协议实验,评估不同音频和文本特征提取器组合及段持续时间,确定30秒段持续时间下MelSpec和ItalianBERT为最优基线。添加DG后,准确率提高2.5%,F1分数提高3.3%,超越现有基准。大量消融研究评估多模态融合、深度架构选择和DG的影响,突出其对稳健且可泛化的抑郁症检测的综合贡献。

英文摘要

Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability. To address this critical issue, we present the first patient-independent multimodal depression detection framework that incorporates domain generalization (DG), jointly leveraging both acoustic and textual modalities. The proposed model integrates bidirectional Long Short-Term Memory (BiLSTM) with intra- and cross-modal attention mechanisms, accompanied by segment-level fusion for decision-making. Generalization is further enhanced by applying a gradient reversal layer inspired by Domain-Adversarial Training of Neural Networks (DANN), which promotes domain-invariant representations by adversarially limiting the model's ability to identify individual speakers, effectively reducing patient-specific bias. Conducting experiments on the Androids-Corpus dataset with a 5-fold cross-validation (CV) protocol, various pairings of audio and text feature extractors were evaluated over different segment durations, determining MelSpec and ItalianBERT as the optimal baseline at a 30-second segment duration. The addition of DG to this baseline yields a 2.5% increase in accuracy and 3.3% in F1-score, achieving 93.2% accuracy, 93.2% precision, 96.2% recall, and 94.2% F1-score, surpassing all existing benchmarks. Extensive ablation studies assess the impact of multimodal fusion, deep architectural choices, and DG, highlighting their combined contribution to robust and generalizable depression detection.

发表机构

  • University of Milan(米兰大学)
  • Gran Sasso Science Institute (GSSI)(大萨索科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑