发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CMDIR,将固定阻抗演示转化为连续变阻抗控制器,作为模仿学习监督,在真实接触任务中提升保留率并降低力波动,FastMPO加速5.8-9.4倍。
AI 中文摘要
CMDIR将流形分解阻抗重定向(MDIR)扩展为将固定阻抗演示转化为连续变阻抗控制器,该控制器也可作为模仿学习的结构化监督。连续任务流形阻抗表示(TMIR)将演化中的任务框架与控制器指令配对。演示相对折衷动力学保留了移动基座传输和控制/物理度量不匹配,从而产生位移、反作用冲量和扰动敏感性准则。质量到快速(Quality-to-Fast)在预定义有限空间内根据开发路径自动编译求解器结构,为每个演示重新实例化该结构,并通过多分辨率评估认证所得候选方案。在三个真实接触任务的225次重定向控制器试验中,与离散MDIR相比,完整CMDIR提高了平均任务代理保留率,并降低了平均姿态偏差、力波动和峰值力。FastMPO相比C-MPO实现了5.8至9.4倍的加速,且闭环结果相当。下游实验证明了完整TMIR监督接口的可学习性;在成功执行中观察到较低的力波动和峰值力,而完成可靠性在不同任务和环境间仍不均匀。
英文摘要
CMDIR extends Manifold-Decomposed Impedance Retargeting (MDIR) to transform fixed-impedance demonstrations into continuous variable-impedance controllers, which can also serve as structured supervision for imitation learning. Continuous Task-Manifold Impedance Representation (TMIR) pairs an evolving task frame with controller instructions. Demo-relative Compromise dynamics retain moving-basis transport and control/physical metric mismatch, yielding displacement, reaction-impulse, and perturbation-sensitivity criteria. Quality-to-Fast automatically compiles a solver structure from development paths within a predefined finite space, re-instantiates that structure for each demonstration, and certifies the resulting candidate by multi-resolution evaluation. Across 225 retargeted-controller trials in three real contact tasks, full CMDIR improves mean task-proxy retention and reduces mean pose deviation, force fluctuation, and peak force relative to discrete MDIR. FastMPO achieves a $5.8$--$9.4\times$ speedup over C-MPO with comparable closed-loop outcomes. Downstream experiments demonstrate learnability of the complete TMIR supervision interface; lower force fluctuation and peak force are observed among successful executions, while completion reliability remains uneven across tasks and environments.
Comments8 pages, 5 figures, 2 tables