EMPIRE:将显式操作规划作为可学习中间表征用于自我中心视角下手运动预测
EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting
浏览论文内容
中文总结 AI 辅助
该研究针对现有手运动预测方法忽略操作过程、梯度干扰的问题,提出两阶段框架EMPIRE,构建EMPIRE-651K数据集,实现了更优的手运动预测精度。
中文摘要 AI 辅助
从自我中心视角的观测中预测灵巧手运动是智能交互系统的基础。现有基于VLM的方法通常直接将观测映射到未来运动,忽略了支配手-物体交互的底层操作过程。此外,端到端优化将操作学习与运动合成耦合,导致运动生成梯度干扰预先学习的感知操作表征。为克服这些局限,我们提出EMPIRE,这是一个两阶段框架,将显式操作规划作为自我中心视角下手运动预测的中间表征。第一阶段:学习规划,EMPIRE首先从多模态上下文学习显式操作规划,以捕获手-物体交互的进展;第二阶段:学习执行,一个运动生成器在冻结规划器表征的条件下合成未来双手运动,防止运动生成梯度影响操作规划。为支撑该方法,我们进一步构建了EMPIRE-651K,这是一个双手运动预测数据集,包含111个任务的650910个训练窗口,每个窗口都配有显式的每只手操作规划。在相同的训练和评估协议下,EMPIRE达到了最先进的预测精度,MPJPE为84.53mm,手指相对误差为38.97mm。我们在该httpsURL发布代码和数据集。
英文摘要
Forecasting dexterous hand motions from egocentric observations is fundamental to intelligent interactive systems. Existing VLM-based methods typically map observations directly to future motions, overlooking the underlying manipulation process that governs hand-object interactions. Moreover, end-to-end optimization couples manipulation learning with motion synthesis, causing motion-generation gradients to interfere with the pre-learned manipulation-aware representations. To overcome these limitations, we propose EMPIRE, a two-stage framework that introduces Explicit Manipulation Planning as an Intermediate Representation for Egocentric hand-motion forecasting. Stage I: Learn to Plan. EMPIRE first learns explicit manipulation plans from multimodal context to capture the progression of hand-object interactions. Stage II: Learn to Act. A motion generator synthesizes future bimanual hand motions conditioned on frozen planner representations, preventing motion-generation gradients from affecting manipulation planning. To support our method, we further construct EMPIRE-651K, a bimanual hand-motion forecasting dataset comprising 650,910 training windows across 111 tasks, each paired with an explicit per-hand manipulation plan. Under identical training and evaluation protocols, EMPIRE achieves state-of-the-art forecasting accuracy, with an MPJPE of 84.53 mm and a finger-relative error of 38.97mm. We release the code and dataset at https://github.com/wangwen-banban/EMPIRE.