InfiniHand:从第一视角视频进行流式世界空间手部运动估计
InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Shanghai Jiao Tong University(上海交通大学)
- University of Science and Technology of China(中国科学技术大学)
- University of Liverpool(利物浦大学)
- Nanyang Technological University(南洋理工大学)
- The University of Hong Kong(香港大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
InfiniHand提出端到端流式前馈框架,联合估计MANO参数、相机轨迹和手部位置,以解决级联方法的误差累积问题,在ARCTIC PA-p上较ViDiHand降低21.4%,并实现11.19 FPS的高吞吐量。
AI中文摘要:
从第一视角(egocentric)视频中进行世界空间手部运动估计,需要在跟踪相机自身运动的同时恢复3D关节手部几何结构。现有方法严重依赖级联的独立手部姿态估计器和SLAM系统,导致误差累积、流程复杂且计算开销巨大。为解决这些局限,我们提出了InfiniHand,一个端到端的流式前馈框架,可直接从未标定的第一视角视频中联合估计MANO参数、相机轨迹和手部位置。InfiniHand将持久时空记忆与以手部为中心的视觉特征相结合,在统一架构中显式耦合相机运动与局部手部几何。我们通过两个渐进阶段训练InfiniHand:首先学习稳健的相机空间手部先验,然后扩展到流式世界空间重建。为支持这一过程,我们汇总了来自多个公开数据集的约5000小时第一视角数据的预训练语料库。大量评估表明,InfiniHand在域内基准上优于最先进的基线,与ViDiHand相比,ARCTIC PA-p降低了21.4%,同时显著减轻了世界空间漂移。此外,InfiniHand对野外视频具有稳健的泛化能力,并以11.19 FPS运行,吞吐量是HaWoR的两倍以上。
英文摘要:
World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhead. To address these limitations, we present InfiniHand, an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video. InfiniHand integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture. We train InfiniHand in two progressive stages by first learning robust camera-space hand priors and then extending to streaming world-space reconstruction. To support this process, we aggregate a pretraining corpus of approximately 5,000 hours of egocentric data across multiple public datasets. Extensive evaluations demonstrate that InfiniHand outperforms state-of-the-art baselines on in-domain benchmarks, achieving a 21.4% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift. Furthermore, InfiniHand generalizes robustly to in-the-wild videos and operates at 11.19 FPS, delivering more than twice the throughput of HaWoR.