发表机构
Department of Information Engineering, University of Padua; Padova Neuroscience Center, University of Padua; Department of Biomedical Sciences, University of Padua; Department of Neuroscience, University of Padua; Information Systems Institute, University of Applied Sciences Western Switzerland (HES-SO Valais)(帕多瓦大学信息工程系; 帕多瓦大学帕多瓦神经科学中心; 帕多瓦大学生物医学科学系; 帕多瓦大学神经科学系; 瑞士西部应用科学大学(HES-SO瓦莱州)信息系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对表面肌电手势识别用于假肢控制时当前架构的局限,提出EMG-CrossFormer,通过级联交叉注意力融合层组合单模态编码器表示,经实验验证该方法能提升仅sEMG解码及多模态融合下的手势识别性能。
AI 中文摘要
通过表面肌电图(sEMG)进行手势识别是假肢控制的基础。深度学习方法已成为该领域的黄金标准,但当前架构难以扩展,随着手部动作数量增加,模型性能通常会下降。卷积操作局部性限制了模型捕捉长程序列模式的能力,单模态设置无法利用来自协调信号的互补信息。本研究引入EMG-CrossFormer,一种用于无缝多模态集成的端到端混合卷积-变压器。它通过级联交叉注意力融合层组合来自任意数量单模态编码器的表示,并使用可学习的手势查询解码融合表示。在四个NinaPro数据集上评估了EMG-CrossFormer,并与六个最先进模型进行基准测试。仅使用sEMG时,在DB2、DB3、DB7和DB10上的平均准确率分别为72.33%、52.48%、79.16%和73.49%。加入惯性信号后性能提高到90.66%、80.40%、92.79%和92.06%。结果表明联合局部-全局特征建模可改善仅sEMG解码,多模态融合显著增强了这一优势。
英文摘要
Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control. In this field, deep learning approaches have become the gold standard. However, current architectures struggle to scale; model performance typically decreases as the number of hand movements increases. Performance degradation is tied to the increased statistical complexity of decoding expanded gesture sets and compounded by the limitations of state-of-the-art methods, which primarily rely on low-latency unimodal convolutional architectures. Convolutions operate locally, limiting model's ability to capture long-range sequential patterns. Unimodal setups cannot leverage complementary information from coordinated signals characterizing movement execution, such as inertial and eye-tracking data. These limitations motivate architectures that integrate local and global features across multimodal physiological sequences. To bridge this gap, this study introduces EMG-CrossFormer, an end-to-end hybrid convolutional-transformer for seamless multimodal integration. EMG-CrossFormer combines representations from an arbitrary number of unimodal encoders through cascaded cross-attention fusion layers, and decodes the fused representations using learnable gesture queries. EMG-CrossFormer was evaluated on four NinaPro datasets (DB2, DB3, DB7, and DB10) and benchmarked against six state-of-the-art models using an increasing number of modalities. Using only sEMG, EMG-CrossFormer achieved mean accuracies of 72.33%, 52.48%, 79.16%, and 73.49% on DB2, DB3, DB7, and DB10, respectively. Incorporating inertial signals improved performance to 90.66%, 80.40%, 92.79%, and 92.06%. These results show that joint local-global feature modeling improves sEMG-only decoding and that multimodal fusion substantially amplifies this benefit, underscoring the value of both design principles for complex hand gesture recognition.
CommentsGitHub repository: see https://github.com/deepPNClab/emg-crossformer