arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16684cs.CV

MEgoVista:野外公制4D手部与头部运动的多视角自我感知估计

MEgoVista: Multi-view Ego-aware Motion Estimation for Metric 4D Hands and Head in the Wild

Jiangong Xiao, Zhihao Zhang, Yifei Dong, Chao Ma, Zhouyi Jin, Zhiwen Hou, Li Liu, Weihuang Chen, Hongbin Sun, Maoqing Yao

首次发表
浏览论文内容

中文总结 AI 辅助

MEgoVista提出一种离线流程,从单个未准备的自我中心视频重建公制双手和头部运动,利用校准立体视觉设定尺度,并在运动捕捉体积内独立验证,为操作学习提供可扩展的标签来源。

中文摘要 AI 辅助

从人类视频中学习操作技能需要以公制单位进行高保真手部运动重建。当今的公制手部标签来自工作室设备和仪器化头戴设备,两者都受限于同样的两个问题:既不能离开预设环境,也没有经过独立参考的校验。不受约束的头戴式记录则提供了相反的权衡,其规模可随佩戴设备的人数扩展。为此,我们提出MEgoVista,一种离线处理流程,可将单个未准备的MEgo View记录转换为统一重力对齐世界坐标系中的公制双手和头部运动。三个特性使其区别于现有的自我中心重建系统:第一,它能在工作室体积和桌面设备无法到达的环境中重建,在检测时确定手部归属,使旁观者的手不进入佩戴者的运动轨迹;第二,其公制基准来自校准立体视觉而非单目先验,在初始化时安装尺度,使策略获得物理单位而非任意坐标;第三,两种输出均在运动捕捉体积内,依据独立的Chingmu光学捕捉进行评分,该协议审计其自身参考,并对方法拒绝预测的部分进行计费。MEgoVista提供了一条从自我中心视频到公制手部监督的可测量路径,拓宽了此类标签可收集的范围。

英文摘要

Learning manipulation from human video requires high-fidelity hand-motion reconstruction in metric units. Today's metric hand labels come from studio rigs and instrumented headsets, and both are confined in the same two ways: neither leaves a prepared setting, and neither is checked against an independent reference. Unconstrained head-worn recording promises the opposite trade-off, scaling with the number of people wearing a device. We therefore introduce MEgoVista, an offline pipeline that turns a single unprepared MEgo View recording into metric two-hand and head motion in one gravity-aligned world frame. Three properties set it apart from existing egocentric reconstruction systems: first, it reconstructs in settings studio volumes and tabletop rigs cannot reach, settling hand ownership at detection so bystander hands stay out of the wearer's trajectory; second, it takes its metric gauge from calibrated stereo rather than a monocular prior, installing scale at initialisation so policies receive physical units, not arbitrary coordinates; third, both outputs are scored inside a motion-capture volume against independent Chingmu optical capture, under a protocol that audits its own reference and charges what a method declines to predict. MEgoVista is offered as a measured route from egocentric video to metric hand supervision, one that widens where such labels can be gathered.

发表机构

  • Northwestern Polytechnical University(西北工业大学)
  • Maniformer

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑