KAD-Net:基于运动学感知解耦学习的单深度图像鲁棒三维手部姿态估计
KAD-Net: Kinematics-Aware Decoupled Learning for Robust 3D Hand Pose Estimation from a Single Depth Image
- Tianjin University(天津大学)
- Chongqing Normal University(重庆师范大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对手部姿态估计中拓扑建模不足与任务干扰问题,提出运动学感知解耦网络KAD-Net,通过FTC模块增强遮挡鲁棒性,并采用任务解耦分层框架,在多个基准上达到最先进精度。
AI中文摘要:
由于手部运动学的复杂性和自遮挡问题,现有的基于单深度图像的三维手部姿态估计方法难以全面建模手部关节之间的拓扑依赖关系。此外,传统的分层多任务架构强制二维关节定位和深度估计共享特征空间,这可能引发相互干扰。为解决这些挑战,我们提出了一种运动学感知解耦学习网络(KAD-Net),用于鲁棒的三维手部姿态估计。具体而言,我们首先设计了一个手指拓扑约束(FTC)模块,以增强远端关节的表示。该模块利用三个连续的手指关节构建局部运动学表示来施加拓扑约束,从而补充远端关节的运动学特征。FTC模块利用可见关节的结构上下文来辅助定位被遮挡的远端关节,从而提升对遮挡的鲁棒性。此外,我们提出了一种任务解耦的分层多任务框架。该框架将二维关节定位与深度估计分离,并为深度回归引入了一种专门的多任务学习策略,有效隔离UV和深度特征,以减轻相互干扰和负迁移。大量实验表明,KAD-Net在多个基准数据集(ICVL、NYU和MSRA)上优于现有方法,在三维手部姿态估计中达到了最先进的精度。KAD-Net的潜在应用包括人机交互、虚拟现实和基于手势的控制系统。
英文摘要:
Due to the complexity of hand kinematics and self-occlusion, existing 3D hand pose estimation methods based on single depth images struggle to comprehensively model the topological dependencies among hand joints. Furthermore, traditional hierarchical multitask architectures enforce a shared feature space for both 2D joint localization and depth estimation, which can induce mutual interference. To address these challenges, we propose a Kinematics-Aware Decoupled Learning Network (KAD-Net) for robust 3D hand pose estimation. Specifically, we first design a Finger Topology Constraint (FTC) module to enhance the representation of distal joints. This module utilizes three consecutive finger joints to construct a local kinematic representation to impose topological constraints, which supplements the kinematic features of the distal joints. The FTC module leverages the structural context from visible joints to assist in locating occluded distal joints, thereby improving robustness to occlusion. Additionally, we propose a task-decoupled hierarchical multitask framework. This framework separates 2D joint localization from depth estimation and incorporates a dedicated multitask learning strategy for depth regression, effectively isolating the UV and depth features to mitigate mutual interference and negative transfer. Extensive experiments demonstrate that KAD-Net outperforms existing methods on several benchmark datasets (ICVL, NYU, and MSRA), achieving state-of-the-art accuracy in 3D hand pose estimation. Potential applications of KAD-Net include human-computer interaction, virtual reality and gesture-based control systems.