一个传感器,全身 - 从单个消费级耳塞IMU恢复3D身体姿态
One Sensor, Whole Body - 3D Body Pose from a Single Consumer Earbud IMU
- University of Cambridge(剑桥大学)
- Chalmers University of Technology(查尔姆斯理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探究单个消费级耳塞IMU能否恢复3D身体姿态,通过构建多模态基准和适配两种循环模型,证明头部IMU可恢复下半身姿态且添加足部IMU无益,指出传感器可靠性而非数量是关键,耳塞为最佳选择。
AI中文摘要:
消费级耳塞已经能够从头部位(人体佩戴最广泛的传感器位置之一)传输惯性运动数据。我们探究单个此类头部IMU能够恢复多少3D身体姿态,以及增加更多消费级传感器是否确实有帮助。我们构建了一个多模态采集流程,同步录制四视角RGB-D视频以及一个AirPods头部IMU和两个Striv鞋垫IMU,事后同步数据流,并使用SAM 3D Body生成伪地面真值,从而构建了一个包含35次采集的单受试者基准数据集,涵盖步态、转身、垂直、日常和临床启发的运动。通过适配两种循环模型家族(IMUPoser和MobilePoser),我们证明单个头部IMU能够以79.0毫米的刚性MPJPE恢复下半身姿态,并以0.809的宏F1分数恢复每只脚的地面接触,且因果变体在流式延迟下保留了大部分精度。在两个模型家族的配对逐次显著性检验中,添加消费级足部IMU从未显著改善姿态估计,并且在四种模型-分割组合中的两种中显著降低了性能;安装偏置探针和仅足部消融实验表明,机制在于鞋垫朝向质量,而非足部放置位置。将输出扩展到20关节全身骨架映射了边界:从头部单独可部分恢复粗略的远端手臂运动,近端上半身姿态则无法恢复,而分阶段微调恢复了朴素联合训练因多任务稀释而牺牲的腿部精度。对于从消费级可穿戴设备学习姿态,传感器可靠性而非传感器数量是此处的约束条件。对于所测试的设备,耳塞是其最佳选择。代码可在以下网址获取:https://this https URL。
英文摘要:
Consumer earbuds already stream inertial motion data from the head, one of the most widely worn sensor locations on the body. We ask how much of the 3D body pose a single such head IMU can recover, and whether adding more consumer sensors actually helps. We build a multimodal capture pipeline that records four-view RGB-D video together with an AirPods head IMU and two Striv insole IMUs, synchronize the streams post-hoc, and generate pseudo-ground-truth with SAM 3D Body, yielding a 35-take single-subject benchmark spanning gait, turning, vertical, everyday, and clinically inspired motions. Adapting two recurrent model families (IMUPoser and MobilePoser), we show that one head IMU recovers lower-body pose at 79.0 mm rigid-MPJPE and per-foot ground contact at 0.809 macro-F1, and that a causal variant retains most of this accuracy at streaming latency. In paired per-take significance tests across both families, adding the consumer foot IMUs never significantly improves pose and significantly degrades it in two of four model-split combinations; a mounting-bias probe and feet-only ablation identify insole orientation quality, not foot placement, as the mechanism. Extending the output to a 20-joint full-body skeleton maps the boundary: gross distal-arm motion is partially recoverable from the head alone, proximal upper-body pose is not, and staged fine-tuning recovers the leg accuracy that naive joint training sacrifices to multi-task dilution. For learned pose from consumer wearables, sensor reliability, not sensor count, is the binding constraint here. For the devices tested, the earbud is its sweet spot. Code is available at https://github.com/ZhilinGuo/one-sensor-whole-body.