发表机构
University of Science and Technology of China; Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences; Arexeni Research and Technologies Inc.; Beijing Sports University; University of Michigan(中国科学技术大学; 中国科学院苏州生物医学工程技术研究所; Arexeni研究与技术公司; 北京体育大学; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MyoMechanix多模态生态系统,构建FKG与CUBIST,建立三类任务,多模态方法提升AQA性能与可解释性,为物理AI提供生物力学基础的技能理解方案。
AI 中文摘要
现有的动作质量评估(AQA)数据集和方法主要依赖RGB、姿态等视觉输入,忽略肌肉力学等生理动态,且常将动作建模为整体模式,这些限制阻碍了基于生物力学的细粒度反馈。本文提出MyoMechanix,一种用于负重动作的多模态生态系统,将运动与肌肉活动对齐。该数据集经专家标注,包含38名受试者完成的20种动作的7500余个样本,配备同步多视角RGB视频、3D姿态、表面肌电(sEMG)及其他生理信号,是目前最大的多模态AQA基准。我们进一步构建健身知识图谱(FKG),将专家标注组织为动作、阶段、关键步骤、错误及纠正反馈间的结构化关系,支持可组合评分与可解释评估。基于这些表示,我们开发CUBIST(可组合本体推理引擎),执行分解-分析-再组合以实现细粒度错误归因与反馈生成。我们还建立了MyoMechanix-AQA、MyoMechanix-VideoQA及新的MyoMechanix-Video2EMG任务。实验表明,多模态感知与结构化表示提升了性能、可解释性与错误归因,CUBIST实现了SOTA结果;VideoQA增强了基于语言的动作理解;Video2EMG为昂贵的肌电传感提供了基于视频的替代方案。MyoMechanix推动技能活动理解向生物力学基础、多模态与可组合推理发展,适用于健身、康复、医疗及机器学习领域的物理AI应用。项目页面:this https URL
英文摘要
Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic patterns. These limitations hinder fine-grained, biomechanically grounded feedback. We introduce MyoMechanix, a multimodal ecosystem for weight-loaded actions that aligns motion with muscle activity. Expert-annotated, it contains 7,500+ samples of 20 actions from 38 subjects, with synchronized multiview RGB video, 3D pose, sEMG, and additional physiological signals, forming the largest multimodal AQA benchmark to date. We further construct the Fitness Knowledge Graph (FKG), which organizes expert annotations into structured relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment. Building on these representations, we develop CUBIST (Compositional Ontological Reasoning Engine), which performs decomposition-analysis-recomposition for fine-grained error attribution and feedback generation. We also establish MyoMechanix-AQA, MyoMechanix-VideoQA, and a novel MyoMechanix-Video2EMG task. Experiments show that multimodal sensing and structured representations improve performance, interpretability, and error attribution, with CUBIST achieving state-of-the-art results; VideoQA enhances language-grounded action understanding; and Video2EMG suggests video-based alternatives to costly EMG sensing. MyoMechanix advances skilled activity understanding toward biomechanically grounded, multimodal, and compositional reasoning for Physical AI applications in fitness, rehabilitation, healthcare, and machine learning. Project page: https://haoyin116.github.io/MyoMechanix/