arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Exo2EgoPose:利用外心演示进行视觉语言引导的自我中心3D手部姿态预测

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting

Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Xiang Li, Hongliang Li

arXiv 2607.15890首次发表:更新:

发表机构

University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言引导的自我中心3D手部姿态预测任务,提出Exo2EgoPose框架,利用外心演示补偿自我中心视角线索不足,含双层外心重建模块和全局到局部调制模块,实验显示该方法优于现有技术,具备人机转移能力且有改进。

AI 中文摘要

从自我中心(Ego)视角感知多模态线索并预测细粒度动作对机器人操作等应用至关重要。以往研究要么主要依赖信息不足的视觉输入预测粗略人体动作,要么遵循VRM/VLA范式,存在机器人数据不足及人机体现差距问题。本文研究视觉语言引导的自我中心3D手部姿态预测(VL-EHPF)任务,提出Exo2EgoPose框架,利用外心(Exo)演示补偿自我中心视角线索。具体包括双层外心重建模块(DERM)和全局到局部调制模块(GLMM)。实验表明该方法优于现有方法,具有有效人机转移能力并在相关数据集上有改进,代码将发布。

英文摘要

Perceiving multimodal cues and forecasting fine-grained actions from an egocentric (Ego) perspective is vital for applications like robot manipulation. However, previous studies either rely mainly on under-informed visual inputs to predict coarse human motions or follow the VRM/VLA paradigm, which suffers from insufficient robot data and the gap between human and robot embodiments. We observe that 3D hand pose naturally serves as a unified representation to bridge human-robot actions. Hence, we investigate an under-explored Vision-Language guided Egocentric 3D Hand Pose Forecasting (VL-EHPF) task, which aims to predict future Ego 3D hand poses from visual observations, a language instruction, and pose states. To overcome the limited field-of-view and highly dynamic motions in the Ego view, we propose a framework dubbed Exo2EgoPose, which innovatively leverages holistic and stable exocentric (Exo) demonstrations as guidance to compensate for partial and dynamic Ego-view cues. Specifically, we introduce a Dual-level Exocentric Reconstruction Module (DERM), which incorporates the paired Exo videos as supervision to reconstruct their video-level and chunked frame-level representations, thereby modeling spatial contexts and temporal dynamics. Then, the Global-to-Local Modulation Module (GLMM) utilizes the reconstructed hierarchical Exo representations for progressive feature refinement via attention mechanisms and adaptive modulation, enabling comprehensive Exo guidance for accurate Ego hand pose forecasting. Extensive experiments on \textit{AssemblyHands}, \textit{Ego-Exo4D}, and our newly constructed \textit{EgoMe-pose} benchmarks show the superiority of our method, which outperforms state-of-the-art methods by a large margin. Moreover, it demonstrates an effective human-to-robot transfer capability and yields improvements on the \textit{CALVIN} dataset.

CommentsAccepted by ACMMM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑