发表机构
Shirley Ryan AbilityLab; Northwestern University(雪莉·瑞安能力实验室; 西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对婴儿无标记三维姿态估计的模型性能问题,采用未标注婴儿视频将Sapiens 2姿态模型跨模型蒸馏到SAM 3D Body模型,提升了该模型在婴儿数据上的二维关键点一致性和三维关节位置误差。
AI 中文摘要
自发运动是婴儿神经运动健康的早期窗口之一,对其评分的结构化临床工具是脑瘫风险的早期有效预测指标。然而这些工具需要受过专门训练的评估者,耗时且存在评估者间的差异,这推动了基于视频的无标记自动评估,尤其是因为基于标记的动作捕捉在婴儿中不实用。但实现无标记捕捉的基础模型几乎完全在成人数据上训练:我们近期的多视角婴儿研究发现,没有单一模型同时表现最佳,不同模型分别具备较强的二维关键点准确率和直接三维身体恢复能力。该研究虽识别出这种权衡,但未解决它。本研究仅使用未标注婴儿视频,将Sapiens 2姿态模型跨模型蒸馏到SAM 3D Body模型中,冻结的教师模型提供密集伪标签,可微渲染器在训练循环中将预测网格与伪标签对齐。在我们前期研究的多视角协议下的11名留存婴儿(18个会话,173条记录)上,微调提升了与Sapiens参考的同视角二维关键点一致性(中位数身体正确关键点百分比@10px从0.22升至0.42,面部从0.22升至0.42),以及普罗克拉斯对齐的三维关节平均位置误差(从25.5降至22.2毫米)。这证明跨模型蒸馏可提升SAM 3D Body模型在婴儿上的性能。
英文摘要
Spontaneous movement is one of the earliest windows onto an infant's neuromotor health, and structured clinical instruments that score it are validated early predictors of cerebral-palsy risk. However, they require specially trained raters, are time-consuming, and carry inter-rater variability. This motivates automated, video-based markerless assessment, especially as marker-based motion capture is impractical in infants. Yet the foundation models that make markerless capture possible are trained almost entirely on adults: our recent multi-view infant study found that no single model is jointly best, with strong 2D keypoint accuracy and direct 3D body recovery split across different models. While that study identifies this trade-off, it does not resolve it. Here, we perform cross-model distillation from the Sapiens 2 pose model into the SAM 3D Body model, using unannotated infant video alone. A frozen teacher supplies dense pseudo-labels, and a differentiable renderer aligns the predicted mesh to them in the training loop. On eleven held-out infants (18 sessions, 173 recordings) under our prior study's multi-view protocol, fine-tuning improves same-view 2D keypoint agreement with the Sapiens reference (median body percentage of correct keypoints @ 10px 0.22 -> 0.42, face 0.22 -> 0.42) and Procrustes-aligned mean per joint 3D position error (25.5 -> 22.2 mm). This demonstrates how cross-model distillation improves SAM 3D Body model performance on infants.
CommentsAccepted to the ECCV 2026 Workshop MoCha