利用二维掩码重建进行三维姿态估计的域适应
Leveraging 2D Masked Reconstruction for Domain Adaptation of 3D Pose Estimation
- UNIST(蔚山科学技术院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种基于掩码图像建模的无监督域适应框架,利用未标记数据并通过前景中心重建和注意力正则化,提升三维姿态估计在跨域场景下的精度,达到最先进水平。
AI中文摘要:
基于RGB的三维姿态估计方法随着深度学习的发展和高质量三维姿态数据集的出现而取得了成功。然而,大多数现有方法在测试图像分布与训练数据分布差异较大时表现不佳。这个问题可以通过在训练中引入多样化数据来缓解,但收集带有相应标签(即三维姿态)的多样化数据并非易事。在本文中,我们提出了一种用于三维姿态估计的无监督域适应框架,该框架通过掩码图像建模(MIM)框架,在利用标记数据的同时也利用未标记数据。我们进一步提出了前景中心重建和注意力正则化,以提高未标记数据使用的有效性。我们在人体和手部姿态估计任务的各种数据集上进行了实验,特别是使用了跨域场景。我们通过在所有数据集上实现最先进的精度,证明了我们方法的有效性。
英文摘要:
RGB-based 3D pose estimation methods have been successful with the development of deep learning and the emergence of high-quality 3D pose datasets. However, most existing methods do not operate well for testing images whose distribution is far from that of training data. However, most existing methods do not operate well for testing images whose distribution is far from that of training data. This problem might be alleviated by involving diverse data during training, however it is non-trivial to collect such diverse data with corresponding labels (i.e. 3D pose). In this paper, we introduced an unsupervised domain adaptation framework for 3D pose estimation that utilizes the unlabeled data in addition to labeled data via masked image modeling (MIM) framework. Foreground-centric reconstruction and attention regularization are further proposed to increase the effectiveness of unlabeled data usage. Experiments are conducted on the various datasets in human and hand pose estimation tasks, especially using the cross-domain scenario. We demonstrated the effectiveness of ours by achieving the state-of-the-art accuracy on all datasets.