发表机构
Institute of Science Tokyo(东京科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人形机器人双臂跟踪中身体招募策略多样且非唯一的问题,提出AMBIT方法,结合CVAE生成与选择器验证,在仿真和实物上显著提升成功率与鲁棒性。
AI 中文摘要
一个具有5自由度手臂的人形机器人仅靠手臂无法跟踪一般的双臂末端执行器轨迹;必须招募骨盆和腰部运动,但具体运动方式及时间并非唯一确定。在固定双脚支撑的Unitree R1上,针对某任务的一组动态有效招募策略(骨盆姿态和腰部轨迹)构成一个多样化的连续流形,基于该流形训练的确定性回归器会进行模式平均,导致策略仅在35%的时间内有效,而条件变分自编码器(CVAE)为52%,16个CVAE样本中最佳者达到82%。我们提出AMBIT:CVAE根据指令轨迹的预览提出策略,一个非学习型选择器对其进行过滤、排序和验证,一个滚动时域循环采用带滞回的方式提交执行。提交的策略与反应式跟踪器运行的同一全身微分逆运动学二次规划(QP)的参考一致,该QP保留对残余误差的控制权。在160个允许有效策略的保留情节中,在扭矩控制器下的完整MuJoCo动力学环境中,AMBIT在3厘米/15度容差下达到85%的成功率,而跟踪器为74%(置信区间不相交),并在48%的情节中在手臂饱和之前招募身体,而跟踪器为35%。由于保持了多样性,训练时未知的约束仅通过选择来强制执行:在五种零样本偏移下,AMBIT在每种偏移上都优于热启动的跟踪器,并匹配了测试时重新优化基线,后者成本高出17倍。在Unitree G1上,超参数不变,该协议重现了有效集的结构,并将与跟踪器的差距扩大到0.85对0.53。五种选定的策略在外部支撑的物理R1上执行,骨盆偏移各不相同,通过编码器正运动学将规划的末端执行器运动跟踪到中位数11毫米,这确立了运动学可实现性,而非平衡性。
英文摘要
A humanoid with 5-DoF arms cannot track generic bimanual end-effector trajectories with its arms alone; pelvis and waist motion must be recruited, but which motion, and when, is not uniquely determined. On a Unitree R1 in fixed double support, the set of dynamically valid recruitment strategies (pelvis pose and waist trajectories) for a task is a diverse continuous manifold, and a deterministic regressor trained on it mode-averages into strategies valid only 35% of the time, against 52% for a conditional variational autoencoder (CVAE) and 82% for the best of 16 CVAE samples. We introduce AMBIT: the CVAE proposes strategies from a preview of the commanded trajectory, a non-learned selector filters, ranks and verifies them, and a receding-horizon loop commits to one with hysteresis. The committed strategy is the reference of the same whole-body differential-IK QP a reactive tracker runs, which keeps authority over residual error. On 160 held-out episodes that admit a valid strategy, in full MuJoCo dynamics under a torque controller, AMBIT reaches 85% success at a 3 cm/15 deg tolerance against 74% for the tracker (disjoint confidence intervals) and recruits the body before the arms saturate in 48% of episodes against 35%. Because diversity is preserved, constraints unknown at training time are enforced by selection alone: under five zero-shot shifts AMBIT beats the warm-started tracker on every shift and matches a test-time re-optimisation baseline 17x more expensive. On a Unitree G1, with hyperparameters unchanged, the protocol reproduces the structure of the valid set and widens the gap over the tracker to 0.85 against 0.53. Five selected strategies execute on the externally supported physical R1, distinct in pelvis excursion and tracking the planned end-effector motion to a median of 11 mm by encoder forward kinematics, which establishes kinematic realisability, not balance.