arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15055eess.SPcs.LG

弥合基于心电图的情绪识别中的差距:深度学习模型的统一评估

Bridging the Gap in ECG-Based Emotion Recognition: A Unified Evaluation of Deep Learning Models

发表机构克拉克森大学 · Affects AI有限责任公司
查看机构详情
  • Clarkson University(克拉克森大学)
  • Affects AI LLC(Affects AI有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

Timothy C Sweeney-Fanelli, Ajan Ahmed, Masudul Imtiaz

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过统一评估框架ARRC和ARDT,整合多个ECG数据集,比较深度学习模型在情绪识别中的泛化性能,建立可复现基准。

中文摘要 AI 辅助

深度学习已催生出众多用于从心电图(ECG)数据中进行自动情绪识别(AER)的架构,但预处理、训练和评估中的不一致性使得直接比较变得困难。大多数研究在均质条件下收集的单个数据集上训练和验证模型,这限制了变异性并引发了对泛化性的担忧。跨数据集验证有时会被使用,但主要评估的是模型适应性而非真正的泛化能力。本研究对AER中著名的深度学习架构进行了比较分析,强调模型泛化性而非数据集适应性。为实现这一基准测试,我们引入了两个开源框架:表征与分类的情感研究(ARRC),一个标准化的基准测试工具包;以及情感研究数据集工具包(ARDT),一个用于跨数据集训练和验证的框架。利用ARDT,我们将三个公开可用的AER数据集——CUADS、ASCERTAIN和DREAMER——整合为一个数据集,增加了传感器类型、记录条件和参与者人口统计特征的变异性。随后,我们使用ARRC通过超参数优化和10折交叉验证评估了三个广泛研究的深度学习模型和两个CNN基线。我们的研究结果提供了关于分类准确性与模型复杂性之间权衡的见解,为AER研究建立了一个可复现的基准。ARRC、ARDT和模型评估的所有源代码均公开可用,以确保透明度并促进进一步研究。

英文摘要

Deep learning has led to numerous proposed architectures for Automated Emotion Recognition (AER) from electrocardiogram (ECG) data, but inconsistencies in preprocessing, training, and evaluation make direct comparisons difficult. Most studies train and validate models on individual datasets collected under homogeneous conditions, limiting variability and raising concerns about generalizability. Cross-dataset validation is sometimes used but primarily assesses model adaptability rather than true generalization. This study presents a comparative analysis of prominent deep learning architectures in AER, emphasizing model generalization over dataset adaptability. To enable this benchmark, we introduce two open-source frameworks: Affective Research on Representations and Classifications (ARRC), a standardized benchmarking toolkit, and Affective Research Dataset Toolkit (ARDT), a framework for inter-dataset training and validation. Using ARDT, we consolidate three publicly available AER datasets, CUADS, ASCERTAIN, and DREAMER, into a single dataset, increasing variability in sensor types, recording conditions, and participant demographics. We then use ARRC to evaluate three widely studied deep learning models and two CNN baselines through hyperparameter optimization and 10-fold cross-validation. Our findings provide insights into the trade-offs between classification accuracy and model complexity, establishing a reproducible benchmark for AER research. All source code for ARRC, ARDT, and model evaluation is publicly available to ensure transparency and facilitate further research.

补充信息

↑