arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于Vision Transformer的相机空间手部姿态精确估计

Estimating Accurate Hand Pose in Camera Space with Vision Transformer

Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia

arXiv 2609.24424首次发表:更新:

发表机构

Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences(中国科学院计算技术研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对单目RGB手部姿态估计中的深度模糊和局部-全局耦合问题,提出变换同构监督与透视信息嵌入,结合帧率感知多数据集训练,在HO3D上CS-MJE指标提升最多37.1%。

AI 中文摘要

基于单目RGB的手部姿态估计已成为计算机视觉领域的一个关键研究前沿。局部手部姿态估计方法预测相对于手腕的手部姿态,而全局手部姿态估计还需要估计手腕在相机坐标系中的位置。然而,这种相机空间估计面临两个基本挑战:(1)单目设置中的深度模糊性,以及(2)手部局部姿态与全局手腕位置在透视投影中的耦合效应。特别地,这种耦合反映了投影由局部手部姿态、手腕位置和相机内参共同决定。为了克服这些挑战,我们的框架提出了两个关键创新:用于手部深度信息提取的变换同构监督(Transformation-Isomorphism Supervision)和用于解决上述局部姿态与手腕位置耦合效应的透视信息嵌入(Perspective Information Embedding),两者都集成在主流的编码器-解码器架构中。此外,我们提出了一种新颖的帧率感知多数据集训练策略,用于序列姿态细化。我们完全集成的方法在HO3D上相对于最先进方法(SOTA)在CS-MJE指标上实现了最多37.1%的提升。项目页面:此https URL。

英文摘要

Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1\% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.

DOI:10.1145/3767308.3836423

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑