发表机构
Fraunhofer HHI; Humboldt University of Berlin(弗劳恩霍夫海因里希·赫兹研究所; 柏林洪堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VR中HMD遮挡上半脸导致情感识别不完整的问题,该研究融合下半脸视频与上半脸EMG,通过晚期融合架构在20名参与者数据集上取得51%宏F1,优于单模态基线,为VR情感识别提供了新方案。
AI 中文摘要
头戴式显示器(HMD)从根本上限制了虚拟现实(VR)中的情感识别:通过遮挡上半脸,使得传统基于图像的面部表情分析不完整,尤其对于需要实时情感评估的应用而言。我们通过融合下半脸视频与被遮挡上半脸的面部肌电信号(EMG)来解决这一挑战,以对7种情感类别(6种基本情感加上中性情感)进行分类。我们引入了一个包含20名参与者的同步多模态数据集,该数据集将下半脸视频与经过验证的情感刺激引发的7通道上半脸EMG配对。在受试者独立测试下,我们提出的晚期融合架构将卷积视觉嵌入与RBF核EMG表示相结合,取得了51%的宏F1值,优于仅使用图像(41%)和仅使用EMG(43%)的基线方法。这些结果表明,在上半脸EMG在HMD引起的视觉遮挡下提供了鲁棒的补充信息,为自然VR环境中的多模态情感识别奠定了基础。该方法有助于实现情感自适应应用,包括沟通训练和治疗干预。该数据集将根据要求在伦理使用协议下共享。
英文摘要
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression analysis incomplete, particularly for applications requiring real-time affective assessment. We address this challenge by fusing lower-face video with facial electromyography (EMG) from the occluded upper face to classify seven emotional categories (six basic emotions plus neutral). We introduce a synchronized multimodal dataset from 20 participants, pairing lower-face video with seven-channel upper-face EMG elicited by validated emotion stimuli. Under subject-independent test, our proposed late-fusion architecture merging convolutional visual embeddings with RBF-kernel EMG representations achieves 51% macro-F1, outperforming both image-only (41%) and EMG-only (43%) baselines. These results demonstrate that upper-face EMG provides robust complementary information under HMD-induced visual occlusion and establish a foundation for multimodal emotion recognition in naturalistic VR environments. This approach facilitates affect-adaptive applications, including communication training and therapeutic interventions. The dataset will be shared upon request under an ethical-use agreement.
CommentsJoint Proceedings of the ACM Intelligent User Interfaces (IUI) Workshops 2026