发表机构
Tsinghua University; Kling Team, Kuaishou Technology; Chinese University of Hong Kong(清华大学; 快手科技 Kling 团队; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出HandsOnWorld框架,通过单目重建从无约束视频中学习手部控制,利用Plücker手部映射解耦相机与手部运动,生成高保真自我中心视频。
AI 中文摘要
我们提出了HandsOnWorld,一个用于手部控制的自我中心视频生成框架,它摒弃了多视角和基于标记的动作捕捉,而是从无约束的单目视频中学习。这种通用性受到可扩展3D手部标注稀缺性的制约:大型自我中心语料库缺乏手指级别的标签,而精确的手部数据集局限于狭窄的仪器化设置,限制了先前手部控制生成器只能处理受限的场景分布。我们通过单目重建直接在野外自我中心视频上标注3D手部,引入了一个以主角为中心的标注流水线,在动作语义、图像质量和3D几何层面过滤重建结果,构建了EgoVid-Pro数据集,该数据集包含103K个片段和大约1200万帧,覆盖多样化的日常场景。为了解决由大自我运动引起的相机-手部纠缠问题,我们进一步提出了Plücker手部映射,这是一种3D感知控制信号,将Plücker射线表示从相机射线扩展到手部表面,在表示层面解耦相机和手部运动。实验表明,我们的方法在重建保真度和控制精度上超越了先前的手部控制生成器,并泛化到先前方法所依赖的实验室数据集之外的分布外日常场景。
英文摘要
We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand annotations from multi-view or marker-based motion capture, confining them to narrow, instrumented scene distributions. To bridge this gap, we introduce a protagonist-centered annotation pipeline that filters monocular 3D reconstructions at the action-semantic, image-quality, and 3D-geometric levels, yielding EgoVid-Pro, a dataset of clean, protagonist-only hand trajectories spanning 103K clips and roughly 12M frames across diverse everyday scenes. These unconstrained scenes exhibit substantial camera ego-motion that is largely absent from tabletop captures, exposing the entanglement of camera and hand motion in existing camera-space control signals. We therefore propose the Plücker Hand Map, which extends Plücker rays from camera geometry to the hand surface, representing hand motion in the same world frame as the camera and disentangling the two motion sources at the representation level. Experiments show that HandsOnWorld outperforms prior methods in visual fidelity and control accuracy and generalizes beyond laboratory settings.
CommentsProject Page: https://shad0wta9.github.io/handsonworld-page/