面向可靠的婴儿姿态估计:一种基于训练动态的噪声标注检测方法
Toward Reliable Infant Pose Estimation: A Training-Dynamics Approach to Noisy Annotation Detection
浏览论文内容
中文总结 AI 辅助
针对早产儿姿态估计中人工标注噪声问题,提出基于训练动态的噪声关键点检测框架,在NeoPose和COCO上分别达到91.9% F1和7.4 AP提升。
中文摘要 AI 辅助
早产儿自发运动分析日益依赖于无标记姿态估计(PE),以直接从视频记录中获取临床相关的运动生物标志物。训练准确的婴儿PE模型需要大量人工标注的关键点,而人工标注 inherently 容易出错。噪声关键点(即相对于其真实解剖位置被错误定位的关键点)在此临床场景中尤其成问题,因为它们可能作为人工伪影传播到重建的关节轨迹中。基于噪声标签学习文献中建立的小损失假设和基于训练动态的样本选择方法,我们提出了一种检测噪声关键点标注的新框架。训练一个混合卷积-注意力模型,从每个关键点的空间坐标和局部视觉特征预测其解剖类别;随后利用由此产生的交叉熵训练动态推导每个关键点的描述符,并通过无监督聚类将其划分为干净子集和噪声子集。我们在NeoPose(一个包含65名住院早产儿的新收集数据集)上,在两种现实噪声场景(随机位置扰动和左右交换)及多种噪声水平下验证了该方法。结果表明,所提方法在噪声关键点检测中达到了高达91.9%的F1分数。该框架进一步泛化到异构的COCO基准,在中等至较高噪声水平下,从训练集中过滤CE检测的噪声关键点在下游姿态估计准确性上带来了可测量的改进(高达7.4个AP点)。
英文摘要
Spontaneous movement analysis in preterm infants relies increasingly on markerless pose estimation (PE) to derive clinically relevant motion biomarkers directly from video recordings. Training accurate infant PE models requires large sets of manually annotated keypoints, and human annotation is inherently prone to error. Noisy keypoints (i.e., keypoints mislocalized with respect to their true anatomical position) are especially problematic in this clinical setting, since they can propagate as artificial artifacts into the reconstructed joint trajectories. Building on the small-loss hypothesis and training-dynamics-based sample selection established in the noisy-label learning literature, we propose a novel framework for detecting noisy keypoint annotations. A hybrid convolutional-attention model is trained to predict the anatomical category of each keypoint from its spatial coordinates and local visual features; the resulting cross-entropy training dynamics are then used to derive per-keypoint descriptors, which are partitioned into clean and noisy subsets via unsupervised clustering. We validate the approach on NeoPose, a newly collected dataset of 65 hospitalized preterm infants, under two realistic noise scenarios (random positional perturbation and left-right swapping) across multiple noise levels. Results show that the proposed approach achieves an F1-score of up to 91.9% in noisy-keypoint detection. The framework further generalizes to the heterogeneous COCO benchmark, where filtering CE-detected noisy keypoints from the training set also yields measurable improvements (up to 7.4 AP points) in downstream pose estimation accuracy at moderate-to-high noise levels.
发表机构
- Università degli Studi “G. d’Annunzio” Chieti-Pescara(基耶蒂-佩斯卡拉“G. 邓南遮”大学)
- KU Leuven(鲁汶大学)
- Università di Teramo(泰拉莫大学)
机构由 AI 辅助整理,请以论文原文为准。