arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04958cs.CVcs.RO

MINT:一种基于可扩展自我中心流水线监督的世界空间相机与手部运动估计统一模型

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

  • ShanghaiTech University(上海科技大学)
  • Wuji Technology(悟吉科技)
  • The University of Hong Kong(香港大学)
  • Zhejiang University(浙江大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Zijie Zhu, Weiren Cai, Yizhou Wang, Zhenjie Yang, Yide Liu, Jiahao Chen, Guanqi He

AI总结:

MINT是首个直接从自我中心RGB视频生成世界空间双手轨迹的统一模型,通过EGOPIPELINE解决标注稀缺问题,在基准测试中精度与速度均提升,可零样本泛化并发布相关资源。

AI中文摘要:

从自我中心视频中恢复世界坐标下的相机与手部运动是活动理解、机器人学习及增强现实的关键能力。现有系统通常将该问题分解为相机运动、深度、手部重建与轨迹优化等独立阶段,导致大量计算开销,且无法对相机与手部运动进行联合建模。我们提出MINT(Minting IN-the-Wild Trajectories),首个能直接从自我中心RGB视频生成完整世界空间双手轨迹的基础模型。MINT从单一共享时空视频表征中,联合预测相机轨迹、相机帧手部状态及逐帧手部存在性,再通过显式坐标变换生成世界空间手部运动。大规模训练该模型颇具挑战,因配对的世界空间相机与手部标注稀缺。为此,我们开发开源标注流水线EGOPIPELINE,可将大量公开自我中心视频转换为结构化相机与手部轨迹监督信号。MINT先在这些大规模伪标签上预训练,再在少量高质量联合标注上微调。在公开基准测试中,MINT在世界空间手部轨迹精度上实现[xxx]的提升,相机轨迹估计上实现[xxx]的提升,端到端轨迹生成速度较标注流水线快[xxx],且能零样本泛化至未见过的自我中心数据集。我们发布该模型、训练与推理代码、标注流水线,以及整理得到的1021小时自我中心轨迹数据集。

英文摘要:

Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity understanding, robot learning, and augmented reality. Existing systems typically decompose this problem into separate stages for camera motion, depth estimation, hand reconstruction, and trajectory refinement, resulting in substantial computational overhead and preventing the joint modeling of camera and hand motion. We introduce MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video. From a single shared spatiotemporal video representation, MINT jointly predicts the camera trajectory, field of view (FoV), camera-frame hand states, and per-frame hand observability, and then produces world-space hand motion via explicit coordinate transformations. Training such a model at scale is challenging, since paired world-space camera and hand annotations are scarce. We therefore develop an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision. MINT is first pretrained on these large-scale pseudo-labels and then fine-tuned on a small set of high-quality camera-and-hand annotations. Across public benchmarks MINT approaches state-of-the-art accuracy without seeing either benchmark in training, reaching 0.945 frame accuracy, 13.646 mm PA-MPJPE-p and 55.058 px EPE-p for camera-frame bimanual reconstruction on HOT3D, 4.690 mm RPE-T and 0.284 degrees RPE-R for camera trajectory, and a 3.67x end-to-end speedup over the labeling pipeline that supervises it. We release the model, training and inference code, labeling pipeline, and a curated 1,021-hour egocentric trajectory dataset.

补充信息

↑