arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35764cs.CVcs.HC

消费级头部与足部IMU的可靠性门控融合用于下肢3D姿态估计

Reliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D Pose

发表机构剑桥大学 · 查尔姆斯理工大学
查看机构详情
  • University of Cambridge(剑桥大学)
  • Chalmers University of Technology(查尔姆斯理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhilin Guo, Boqiao Zhang, Oszkár Urbán, Josef Bengtson, Hakan Aktas, Wenzhao Li, Siyu Hong, Kyle Fogarty, Chenliang Zhou, Ali Senguel, Cengiz Oztireli

首次发表
浏览论文内容

中文总结 AI 辅助

针对消费级IMU不可靠问题,提出通道级可靠性门控融合模型,在稀疏惯性下肢姿态估计中显著提升精度,优于静态融合和未门控方案,并有效应对传感器故障。

中文摘要 AI 辅助

稀疏惯性姿态估计有望实现无需摄像头的消费级设备运动捕捉,但消费级传感器并不可靠:固件融合的朝向存在偏差,安装方式在不同会话间存在差异,且数据流会漂移或丢失。在一个新的包含35次采集的单被试基准测试中,我们将耳塞式头部惯性测量单元(IMU)与两个智能鞋垫足部IMU配对(使用SAM-3D-Body伪地面真值标签),证明了可靠性问题是通道级别的:通道消融实验将足部加速度识别为信息量最大的输入(66.6毫米对比仅头部时的79.0毫米),而固件融合的足部朝向则是破坏增益的短板。因此,我们让模型学习对每个流的每个通道的信任程度:每个流每个通道块设置一个时间门控,并使用在合成损坏的预训练数据上训练的辅助可靠性目标进行训练。通道门控模型在干净数据上是我们学习的融合方案中最准确的(69.4毫米对比静态融合的83.7毫米和未门控的86.6毫米),并且在每种模拟故障下(训练中的偏差;仅评估时的漂移和丢失)均表现最佳;其门控在无测试时监督的情况下抑制了干净真实数据上固有偏差的足部朝向通道,并以0.92-0.999的AUROC标记丢失突发。两个对比实验:预先知道会失败的通道丢弃在足部故障下表现平稳,但在意外流失败时崩溃(头部丢失:92.9毫米对比79.3毫米);微调的HMD-Poser在干净数据上更准确(64.4毫米)且在漂移下名义上更优,但在偏差或丢失下无显著配对差异,但最坏情况下的退化更大(+16.1毫米对比+3.5毫米,单种子)。学习门控可靠性而非传感器数量是部署稀疏惯性捕捉的关键。代码可在以下网址获取:https URL。

英文摘要

Sparse inertial pose estimation promises camera-free motion capture from consumer devices, but consumer sensors are unreliable: firmware-fused orientations are biased, mounting varies between sessions, and streams drift or drop out. On a new 35-take single-subject benchmark pairing an earbud head inertial measurement unit (IMU) with two smart-insole foot IMUs (SAM-3D-Body pseudo-ground-truth labels), we show the reliability problem is channel-level: a channel ablation isolates foot acceleration as the most informative input (66.6 mm vs. 79.0 mm head-only) and the firmware-fused foot orientation as the liability that destroys the gain. We therefore let the model learn how much to trust each channel of each stream: one temporal gate per stream per channel block, trained with an auxiliary reliability objective on synthetically corrupted pretraining data. The channel-gated model is the most accurate of our learned fusion arms on clean data (69.4 mm vs. 83.7 static, 86.6 ungated) and under every simulated fault (bias in training; drift, dropout eval-only); its gates suppress the natively biased foot-orientation channels on clean real data without test-time supervision and flag dropout bursts at 0.92-0.999 AUROC. Two contrasts: dropping a channel known a priori to fail is flat across foot faults but collapses when an unanticipated stream fails (head dropout: 92.9 vs. 79.3 mm); and a fine-tuned HMD-Poser is more accurate on clean data (64.4 mm) and nominally under drift, with no significant paired difference under bias or dropout, but a larger worst-case degradation from clean (+16.1 vs. +3.5 mm, single seed). Learning to gate reliability instead of sensor count is the lever for deployable sparse inertial capture. Code is available at https://github.com/ZhilinGuo/reliability-gated-imu-fusion.

补充信息

↑