EVFormer:用于双手手部姿态估计的以自我为中心的视觉-肌电双向注意力模型
EVFormer: An Egocentric Vision-EMG Bidirectional Attention Model for Bimanual Hand Pose Estimation
AI总结:
EVFormer通过双向交叉注意力融合自我中心视觉与前200毫秒双侧腕部sEMG信号,估计44个手指和腕部关节角度,在296个测试样本上实现11.482度平均绝对误差,较视觉基线降低13.20%。
AI中文摘要:
以自我为中心的双手手部姿态估计对于虚拟交互、可穿戴控制和康复非常重要,但视觉观察常常因自遮挡、手-手接触和物体操作而退化。我们提出了EVFormer,一种多模态框架,将当前RGB帧与前200毫秒的双侧腕部表面肌电(sEMG)信号相结合,以估计44个手指和腕部关节角度。EVFormer分别编码视觉空间特征和sEMG时间特征,通过顺序双向交叉注意力实现跨模态信息交换,并使用特征级门控融合整合两种模态。我们在单参与者可行性研究中评估了EVFormer,使用一个同步的公共EgoEMG记录,并按时间顺序划分训练、验证和测试集。在296个测试样本上,EVFormer实现了11.482度的平均绝对误差,而仅视觉、仅sEMG、晚期融合和训练均值基线分别为13.228-13.610度。这相当于与仅视觉模型相比相对误差减少13.20%,与晚期融合相比减少14.23%。EVFormer在五个评估手势类别中的四个也实现了最低误差。这些结果初步证明,以自我为中心的视觉与sEMG之间的特征级交互可以改善双手手部姿态估计。需要在不同参与者、记录会话、传感器放置和真实世界交互条件下进行进一步评估,以确定该方法的泛化性。
英文摘要:
Egocentric bimanual hand pose estimation is important for virtual interaction, wearable control, and rehabilitation, but visual observations are often degraded by self-occlusion, hand-hand contact, and object manipulation. We propose EVFormer, a multimodal framework that combines the current RGB frame with the preceding 200 ms of bilateral wrist surface electromyography (sEMG) to estimate 44 finger and wrist joint angles. EVFormer separately encodes visual spatial features and sEMG temporal features, enables cross-modal information exchange through sequential bidirectional cross-attention, and integrates the two modalities using feature-wise gated fusion. We evaluate EVFormer in a single-participant feasibility study using one synchronized public EgoEMG recording with chronologically separated training, validation, and test splits. On 296 test samples, EVFormer achieves a mean absolute error of 11.482 degrees, compared with 13.228-13.610 degrees for vision-only, sEMG-only, late-fusion, and training-mean baselines. This corresponds to relative error reductions of 13.20% compared with the vision-only model and 14.23% compared with late fusion. EVFormer also achieves the lowest error in four of the five evaluated gesture classes. These results provide preliminary evidence that feature-level interaction between egocentric vision and sEMG can improve bimanual hand pose estimation. Further evaluation across participants, recording sessions, sensor placements, and real-world interaction conditions is required to establish the generalizability of the approach.