发表机构
University of Texas at San Antonio; Kennesaw State University(德克萨斯大学圣安东尼奥分校; 肯尼索州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出用MI-CNN模型结合多模态数据与SHAP可解释性分析,实现VR中平衡与不平衡姿态的二分类,准确率达96.76%,降维后仍保持良好性能,为安全VR系统提供支撑。
AI 中文摘要
确保安全的虚拟现实(VR)体验需要能够预测并响应用户失衡状态的系统。尽管已有研究探讨了跌倒预测和运动病,但多数方法基于回归,而姿态状态分类仍较少被探索。本研究比较了用于在视觉干扰下对VR中姿态状态进行分类的机器学习(ML)和深度学习(DL)模型。我们使用了包含运动学、肌电图(EMG)和皮肤电活动(EDA)信号的多模态数据集。数据被处理为区分平衡与不平衡姿态状态的二分类任务,且采用按参与者降采样的方法解决类别不平衡问题。所有模型均采用留一参与者交叉验证(LOPO)进行评估,以测试对未见过参与者的泛化能力。在所有模型中,受Mamba启发的卷积神经网络(MI-CNN)达到了96.76%的最高准确率。SHapley加性解释(SHAP)分析提升了可解释性,并识别出最具影响力的分类因素。SHAP结果显示运动学特征占主导,表明身体运动模式对检测VR中的不平衡具有信息价值。我们还使用按SHAP重要性排名的前三分之二特征评估了MI-CNN。尽管输入维度减少了33%,模型仍保持性能,达到0.957的准确率和0.957的F1分数,与全特征模型相比仅下降约1%。这些发现表明,多模态感知、时序深度学习和可解释人工智能可支持VR中与平衡相关不稳定状态的可靠分类。准确识别不平衡姿态状态可提高对跌倒风险的感知,并指导更安全、自适应的VR系统,该系统可在不稳定时做出响应,同时提升用户安全与体验。代码可从该https URL获取。
英文摘要
Ensuring a safe virtual reality (VR) experience requires systems that can predict and respond when users lose their balance. Although prior work has examined fall prediction and motion sickness, many approaches are regression-based and postural state classification remains less explored. This study compares machine learning (ML) and deep learning (DL) models for classifying postural states in VR under visual perturbations. We used a multimodal dataset containing kinematic, electromyographic (EMG), and electrodermal activity (EDA) signals. The data were prepared for a binary task to distinguish balanced from imbalanced postural states, and participant-wise downsampling addressed class imbalance. All models were evaluated with Leave-One-Participant-Out (LOPO) cross-validation to test generalization to unseen participants. Among the models, the Mamba-inspired CNN (MI-CNN) achieved the highest accuracy of 96.76%. SHapley Additive exPlanations (SHAP) analysis improved interpretability and identified the most influential classification factors. The SHAP results showed that kinematic features were dominant, indicating that body-motion patterns are informative for detecting imbalance in VR. We also evaluated MI-CNN using only the top two-thirds of features ranked by SHAP importance. Despite a 33% reduction in input dimensionality, the model maintained performance, achieving 0.957 accuracy and 0.957 F1-score, with about a 1% decrease compared with the full-feature model. These findings suggest that multimodal sensing, temporal deep learning, and explainable AI can support reliable classification of balance-related instability in VR. Accurate recognition of imbalanced postural states may raise awareness of fall risk and guide safer, adaptive VR systems that respond to instability while improving user safety and experience. Code is available at: https://github.com/NipaAnjum/MI-CNN.
CommentsAccepted in The 25th IEEE International Symposium on Mixed and Augmented Reality (ISMAR)