发表机构
Sanofi Digital R&D(赛诺菲数字研发)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出STA-TFM,一种基于Transformer的多视角姿态估计架构,结合单目特征提取器与融合Transformer,并利用数据生成流程,在多个数据集上显著降低姿态误差,且能处理噪声和缺失输入。
AI 中文摘要
单目三维人体姿态估计(HPE)由于深度模糊、遮挡以及时间一致性的需求仍然具有挑战性。虽然多视角方法在精度上优于单目方法,但它们通常需要复杂的设置。我们引入了STA-TFM,一种基于Transformer的架构,结合空间和时间信息进行多视角姿态估计。该方法利用DSTformer(一种单目特征提取器)来捕获每个视角内的长距离姿态依赖关系。然后,一个融合Transformer跨视角聚合信息,以生成连贯的三维估计。为解决训练数据稀缺问题,我们使用一个数据生成流程,可将任何现有的三维姿态数据集转换为具有可控参数的多视角设置。在多个数据集上的实验表明,STA-TFM优于现有的无需相机参数的多视角方法。STA-TFM在DHP19数据集上实现了平均每关节位置误差(MPJPE)和平均每关节速度误差(MPJVE)分别降低50.9%和49.5%。此外,它在HAA4D上分别实现了6.7%和7.7%的降低,在TotalCapture上实现了15.2%的MPJPE降低。STA-TFM能够处理有噪声和缺失的二维输入,支持在医疗健康监测、运动评估和沉浸式技术中的潜在部署。代码、训练检查点和数据可在以下网址获取:此HTTPS URL。
英文摘要
Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-view methods provide superior accuracy over monocular approaches, they often require complex setups. We introduce STA-TFM, a transformer-based architecture that combines spatial and temporal information for multi-view pose estimation. The approach leverages DSTformer, a monocular feature extractor, to capture long-range pose dependencies within each view. A fusion transformer then aggregates information across views to produce coherent 3D estimates. To address training data scarcity, we use a data generation pipeline that transforms any existing 3D pose dataset into multi-view setups with controllable parameters. Experiments on various datasets demonstrate that STA-TFM outperforms existing camera-parameter-free multi-view methods. STA-TFM achieves 50.9% and 49.5% reductions in mean per joint position error (MPJPE) and mean per joint velocity error (MPJVE) on the DHP19 dataset. Furthermore, it achieves 6.7% and 7.7% respective reductions on HAA4D, and a 15.2% MPJPE reduction on TotalCapture. STA-TFM handles noisy and missing 2D inputs, supporting potential deployment in healthcare monitoring, athletic assessment, and immersive technologies. Code, training checkpoints, and data are available at https://zenodo.org/records/22832620.