arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27595cs.HC

ViMoWear:视觉运动引导的sEMG-IMU表示学习用于受试者无关的拇指手势识别

ViMoWear: Visual Motion-Guided sEMG-IMU Representation Learning for Subject-Independent Thumb Gesture Recognition

发表机构爱丁堡大学
查看机构详情
  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenjuan Zhong, Chenfei Ma, Kianoush Nazarpour

首次发表
浏览论文内容

中文总结 AI 辅助

ViMoWear通过视觉运动引导的对比学习和掩蔽重建,仅用可穿戴信号实现受试者无关的拇指手势识别,提升泛化性能。

中文摘要 AI 辅助

可穿戴传感使得人机交互、增强现实和假肢控制中的直观手势识别成为可能,然而受试者无关的识别仍然具有挑战性,因为可穿戴信号仅提供对手部运动的间接且高度受试者特定的观察。尽管视觉信息可以改善可穿戴手势识别,但在推理时要求视觉信息会增加传感复杂性并限制实际部署。我们提出了ViMoWear,一个视觉运动引导的框架,利用同步的3D手部运动作为仅训练时的监督,而在推理时仅需可穿戴传感进行手势分类。具体而言,运动引导的跨受试者对比学习(MGCL)促进受试者鲁棒的表示,而拇指感知的掩蔽运动重建(TMMR)保留细粒度的运动信息。在同步的sEMG-IMU-姿态数据集上的留一受试者实验表明,在多种传感配置下,相对于监督基线有一致的改进,而学习到的表示也支持无分类器的检索。所提出的仅训练时的视觉运动监督提高了可穿戴表示对未见受试者的泛化能力。

英文摘要

Wearable sensing enables intuitive hand gesture recognition for human--computer interaction, augmented reality, and prosthetic control, yet subject--independent recognition remains challenging because wearable signals provide only indirect and highly subject-specific observations of hand motion. Although visual information can improve wearable gesture recognition, requiring it during inference increases sensing complexity and limits practical deployment. We propose ViMoWear, a visual-motion-guided framework that leverages synchronized 3D hand motion as training-only supervision while requiring only wearable sensing for gesture classification at inference. Specifically, Motion-Guided Cross-Subject Contrastive Learning (MGCL) promotes subject-robust representations, and Thumb-Aware Masked Motion Reconstruction (TMMR) preserves fine-grained motion information. The leave-one-subject-out experiments on a synchronized sEMG--IMU--pose dataset demonstrate consistent improvements over supervised baselines across multiple sensing configurations, while the learned representations also support classifier-free retrieval. The proposed training-only visual motion supervision improves the generalization of wearable representations to unseen subjects.

↑