发表机构
Korea Electronics Technology Institute; Seoul National University; Korea University(韩国电子技术研究院; 首尔大学; 高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单段运动片段训练的人形机器人移动操作策略仅复现终点运输的问题,提出距离条件化参考重组(DCRR),通过重组终止段并蒸馏距离条件化监督,在四种任务中显著降低距离误差,并支持硬件上的距离调制。
AI 中文摘要
运动跟踪可以从单个重定向的运动片段中复现人形机器人的移动操作,但基于固定参考训练的策略主要复现其演示的运输结果。尽管源轨迹经过了中间物体位移,但运输终止仅在终点得到演示。我们将这一不匹配识别为终止与通过之间的差距:中间位移被观察为通过状态,而非终止完成的结果。我们引入了距离条件化参考重组(DCRR),该方法将演示的终止段重新定位到中间运输状态。一个冻结的跟踪教师策略在闭环动力学下重放重组后的参考,保留的轨迹根据其实现的物体放置位置进行重新标注,并蒸馏到一个无参考策略中。该过程从源运动中编码的交互行为构建了距离条件化的监督。在搬运、踢推、蹲推和拖拽任务中,DCRR-BC产生了与命令相关的运输,整体归一化距离平均绝对误差(MAE)为0.15,而仅使用源数据的行为克隆为0.28。强化学习微调进一步提高了训练模拟器和模拟到模拟迁移中的命令响应和执行鲁棒性。最后,硬件实验展示了所有四种交互模式下的运输距离调制。
英文摘要
Motion tracking can reproduce humanoid loco-manipulation from a single retargeted motion clip, but a policy trained on a fixed reference primarily reproduces its demonstrated transport outcome. Although the source trajectory visits intermediate object displacements, transport termination is demonstrated only at its endpoint. We identify this mismatch as the termination-versus-passage gap: intermediate displacements are observed as passage states rather than termination-complete outcomes. We introduce Distance-Conditioned Reference Recomposition (DCRR), which relocates the demonstrated termination segment to intermediate transport states. A frozen tracking teacher replays the recomposed references under closed-loop dynamics, and the retained trajectories are relabeled by their achieved object placements and distilled into a reference-free policy. This procedure constructs distance-conditioned supervision from the interaction behavior encoded in the source motion. Across Carry, Kick-Push, Crouch-Push, and Drag, DCRR-BC produces command-dependent transport with an overall normalized distance mean absolute error (MAE) of 0.15, compared with 0.28 for source-only behavior cloning. RL fine-tuning further improves the command response and execution robustness in the training simulator and under sim-to-sim transfer. Finally, hardware experiments demonstrate transport-distance modulation across all four interaction modes.