arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从像素学习人体关节力矩

Learning human joint torques from pixels

Chen Chen, Rui Cheng

arXiv 2608.09083首次发表:更新:

发表机构

Beijing Normal-Hong Kong Baptist University; AnHui University(北京师范大学-香港浸会大学联合国际学院; 安徽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于视觉的逆动力学数据集VID及基准,构建VID-Network模型,在VID数据集上实现关节力矩估计误差显著降低,为生物力学推理提供实用基础。

AI 中文摘要

从视觉观测中估计人体关节力矩是将生物力学分析从受控实验室推向真实运动场景的关键步骤。现有力矩估计方法通常依赖表面肌电图、运动捕捉标记、测力台或模拟模仿数据,这限制了它们在普通RGB图像上的适用性。在本研究中,我们引入VID,这是一个基于视觉的逆动力学数据集和基准,用于直接从真实单目图像预测人体关节力矩。VID包含63369个同步帧,带有真实人体图像、运动学标注、人体测量属性以及OpenSim衍生的动态标签,为真实图像的逆动力学提供配对的视觉和生物力学监督。我们还定义了标准化评估协议,涵盖整体力矩估计、关节特定分析和动作特定预测。为建立强大的参考模型,我们提出VID-Network,它结合了经过姿态预训练的空间概率特征、标记回归和时序力矩推理,以从图像序列中恢复关节力矩。在VID上的实验表明,VID-Network实现了1.7612 N·m/kg的整体平均逐关节误差(mPJE),比最佳对比基线提升了39.81%,且在所有评估的关节类型和大多数动作类别中都获得了最低误差。VID为视觉驱动的人体逆动力学建立了首个实用基准,并为在约束较少的环境中研究生物力学推理提供了基础。

英文摘要

Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-world movement scenarios. Existing torque estimation methods typically depend on surface electromyography, motion-capture markers, force plates, or simulated imitation data, which limits their applicability to ordinary RGB images. In this work, we introduce VID, a vision-based inverse dynamics dataset and benchmark for predicting human joint torques directly from real monocular images. VID contains 63,369 synchronized frames with real human images, kinematic annotations, anthropometric attributes, and OpenSim-derived dynamic labels, providing paired visual and biomechanical supervision for real-image inverse dynamics. We further define a standardized evaluation protocol covering overall torque estimation, joint-specific analysis, and action-specific prediction. To establish a strong reference model, we propose VID-Network, which combines pose-pretrained spatial probabilistic features, marker regression, and temporal torque inference to recover joint torques from image sequences. Experiments on VID show that VID-Network achieves an overall mPJE of 1.7612 N$\cdot$m/kg, improving over the best compared baseline by 39.81\%, and obtains the lowest error across all evaluated joint types and most action categories. VID establishes a first practical benchmark for vision-driven human inverse dynamics and provides a foundation for studying biomechanical inference in less constrained environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑