arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HandMvNet:基于多视图交叉注意力融合的实时3D手部姿态估计方法

HandMvNet: Real-Time 3D Hand Pose Estimation Using Multi-View Cross-Attention Fusion

Muhammad Asad Ali, Nadia Robertini, Didier Stricker

arXiv 2608.20093首次发表:更新:

发表机构

German Research Center for Artificial Intelligence (DFKI); University of Kaiserslautern-Landau (RPTU)(德国人工智能研究中心; 凯撒斯劳滕-兰道大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HandMvNet是无需相机参数输入、推理时间短且性能优异的实时3D手部姿态估计方法,适用于多视图图像的实时应用。

AI 中文摘要

本研究提出HandMvNet,是首批专为从多视图相机图像估计3D手部运动与形状设计的实时方法之一。与存在尺度-深度模糊问题的单目方法不同,该方法通过多视图注意力融合机制有效整合多视角特征,确保获得一致且准确的绝对手部姿态与形状。与以往多视图方法相比,本方法无需将相机参数作为输入即可学习3D几何。HandMvNet大幅缩短了推理时间,同时达到与现有最优方法相当的结果,适用于实时应用。在公开数据集上的评估显示,HandMvNet在相同设置下定性与定量均优于此前方法,代码可在指定网址获取。

英文摘要

In this work, we present HandMvNet, one of the first real-time method designed to estimate 3D hand motion and shape from multi-view camera images. Unlike previous monocular approaches, which suffer from scale-depth ambiguities, our method ensures consistent and accurate absolute hand poses and shapes. This is achieved through a multi-view attention-fusion mechanism that effectively integrates features from multiple viewpoints. In contrast to previous multi-view methods, our approach eliminates the need for camera parameters as input to learn 3D geometry. HandMvNet also achieves a substantial reduction in inference time while delivering competitive results compared to the state-of-the-art methods, making it suitable for real-time applications. Evaluated on publicly available datasets, HandMvNet qualitatively and quantitatively outperforms previous methods under identical settings. Code is available at github.com/pyxploiter/handmvnet.

CommentsPublished at VISAPP 2025. 8 pages, 7 figures

Journal refProceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 2: VISAPP (2025), pp. 555-562

DOI:10.5220/0013107300003912

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑