发表机构
Meta Reality Labs Research; UC Davis; UNC Chapel Hill; UC Berkeley; MIT; CMU(Meta现实实验室; 加州大学戴维斯分校; 北卡罗来纳大学教堂山分校; 加州大学伯克利分校; 麻省理工学院; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对人类演示重定向至机器人时的动力学不可行等问题,提出生成式神经重定向(GNR)方法,样本效率远高于MPC,构建了大规模灵巧操作数据集。
AI 中文摘要
人类演示是学习灵巧操作的可扩展数据源,但实体差距导致人类动作无法直接在机器人上执行。逆运动学(IK)可高效将人类动作重定向至机器人,但忽略动力学,常产生不可行动作。强化学习(RL)和基于采样的模型预测控制(MPC)常用于生成动力学可行的动作,但二者均样本效率低且对超参数敏感:RL存在训练成本高、不稳定及奖励工程繁琐的问题;MPC虽避免策略优化,但对每条轨迹单独重定向,求解一条不会降低下一条的难度,且采样成本随数据集规模和任务难度快速增长。我们假设动力学可行的轨迹集中在演示间共享的低维流形附近,因此重定向可简化为在人类动作条件下从该流形采样,而非为每个演示求解新的优化问题。我们提出生成式神经重定向(Generative Neural Retargeting, GNR),其使用流匹配模型采样可行轨迹。GNR的样本量仅为MPC的8.5%,性能优于MPC,成功率达56.20%,而MPC为27.20%。GNR可用于大规模、长 horizon、毫米级精度人类演示的可扩展高效重定向:通过在真实到仿真数据引擎中应用GNR,我们生成了包含密集接触力标签的灵巧操作数据集,涵盖22.3万次演示和3300种物体几何形状。
英文摘要
Human demonstrations are a scalable data source for learning dexterous manipulation, but the embodiment gap prevents human motion from being executed directly on robots. Inverse kinematics (IK) retargets human motion to robots efficiently but ignores dynamics, often producing infeasible motions. Reinforcement learning (RL) and sampling-based model predictive control (MPC) are commonly employed to yield dynamically feasible motions, but both are sample-inefficient and sensitive to hyperparameters. RL suffers from costly and unstable training and tedious reward engineering; MPC avoids policy optimization, yet retargets each trajectory in isolation, and solving one does not make the next easier. Sampling cost grows rapidly with dataset size and task difficulty. We hypothesize that dynamically feasible trajectories concentrate near a low-dimensional manifold shared across demonstrations, so that retargeting can be reduced to sampling from that manifold, conditioned on human motion, rather than solving a fresh optimization problem for every demonstration. We propose \textbf{Generative Neural Retargeting} (GNR), which uses a flow matching model to sample feasible trajectories. GNR outperforms MPC with only $8.5\%$ of the samples required by MPC, achieving a success rate of $56.20\%$ compared to $27.20\%$ for MPC. GNR can be used for scalable and efficient retargeting of large-scale, long-horizon, and millimeter precision human demonstrations: by applying GNR within a real-to-sim data engine, we produce a dexterous manipulation dataset with dense contact-force labels, spanning $223$k demonstrations and $3.3$k object geometries.