arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20991cs.LG

从压力到情感:跨可穿戴传感器模态的生理情感识别多模态深度学习

From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

  • Howard University(霍华德大学)

机构由 AI 辅助整理,请以论文原文为准。

Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge

AI总结:

本研究比较了LSTM、TCN和Transformer在WESAD和EmoWear数据集上的生理情感识别性能,发现架构优劣依赖数据集,多模态传感优于单模态,4 Hz采样频率为实用选择。

AI中文摘要:

利用可穿戴传感器进行生理情感识别在心理健康监测、情感计算和人机交互中具有重要应用。然而,现有研究通常评估单一模型、传感配置或数据集,限制了我们对这些因素如何影响识别性能的理解。我们提出了一项关于生理情感识别的时序深度学习架构的比较研究,使用两个多模态可穿戴数据集:WESAD和EmoWear。在仅腕部、仅胸部和多模态传感配置下,使用参与者独立的留一受试者交叉验证(LOSO-CV)评估了双向长短期记忆网络(LSTM)、时序卷积网络(TCN)和Transformer模型。我们还研究了软投票集成、传感器消融、采样频率和基于梯度的显著性。Transformer在WESAD上达到了最高的多模态准确率(99.02% ± 0.51%),而LSTM在EmoWear上对唤醒度(91.80% ± 1.06%)和效价(89.96% ± 0.36%)均取得了最佳多模态准确率。这些结果表明,相对架构性能取决于数据集特征,而非某一种架构普遍更优。多模态传感在两个数据集上均持续优于仅腕部和仅胸部配置。采样频率分析显示,4 Hz提供了一个实用的工作点,其性能与更高频率相当,但训练成本大幅降低。这些发现为可穿戴生理情感识别的架构、传感模态和采样频率选择提供了指导。

英文摘要:

Physiological emotion recognition using wearable sensors has important applications in mental health monitoring, affective computing, and human-computer interaction. However, existing studies typically evaluate a single model, sensing configuration, or dataset, limiting our understanding of how these factors influence recognition performance. We present a comparative study of temporal deep learning architectures for physiological emotion recognition using two multimodal wearable datasets: WESAD and EmoWear. Bidirectional long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer models are evaluated under wrist-only, chest-only, and multimodal sensing configurations using participant-independent leave-one-subject-out cross-validation (LOSO-CV). We also investigate soft-voting ensembles, sensor ablation, sampling frequency, and gradient-based saliency. The Transformer achieved the highest multimodal accuracy on WESAD (99.02% +/- 0.51%), whereas the LSTM achieved the best multimodal accuracy on EmoWear for both arousal (91.80% +/- 1.06%) and valence (89.96% +/- 0.36%). These results show that relative architecture performance depends on dataset characteristics rather than one architecture being uniformly superior. Multimodal sensing consistently outperformed wrist-only and chest-only configurations across both datasets. Sampling-frequency analysis showed that 4 Hz provides a practical operating point, with performance comparable to higher frequencies at substantially lower training cost. These findings provide guidance for selecting architectures, sensing modalities, and sampling frequencies for wearable physiological emotion recognition.

补充信息

↑