发表机构
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对仿真到真实系统辨识中轨迹无法区分参数导致估计不可靠的问题,提出IDTD框架,通过舒尔补分数设计互补轨迹区间,在多仿真环境及真实K1人形机器人上提升了参数辨识精度与策略迁移性能。
AI 中文摘要
基于采样的系统辨识通过调整模拟器以复现目标系统动力学来估计具有物理意义的参数,为提升仿真到真实(sim-to-real)迁移提供了可解释的方法。然而,当采集的轨迹无法区分不同参数的影响时,多个参数组合可复现这些轨迹,导致参数估计不可靠。为解决该挑战,我们提出信息解耦轨迹设计框架(Informationally Decoupled Trajectory Design, IDTD),该框架基于费舍尔信息矩阵推导的舒尔补分数构建探索策略的目标。为在探索目标中忠实地反映参数可分性,IDTD针对每个参数的信息对分数进行归一化,为每个参数选择最有利的轨迹段,并对数聚合所得分数。因此,优化后的轨迹由互补区间组成,每个区间暴露一组不同的参数,其在该区间内对运动的贡献可明确归因。在从线性动力学系统到Go2四足机器人、G1人形机器人及Crazyflie四旋翼的多种仿真环境中,IDTD相较于最强的现有主动探索基线,平均降低了39.6%的参数辨识误差,并实现了更优的下游策略迁移。此外,我们在真实K1人形机器人上验证了所提出的轨迹设计,证明辨识得到的参数可准确捕捉真实系统的动力学特性。
英文摘要
Sampling-based system identification estimates physically meaningful parameters by tuning a simulator to reproduce the target system dynamics, providing an interpretable approach to improving sim-to-real transfer. Yet when the collected trajectories do not distinguish the effects of different parameters, multiple parameter combinations can reproduce those trajectories, leading to unreliable parameter estimates. To address this challenge, we introduce an Informationally Decoupled Trajectory Design framework (IDTD), which formulates the objective for the exploration policy built on the Schur complement score derived from the Fisher information matrix. To faithfully reflect parameter separability in the exploration objective, IDTD normalizes the score against the per-parameter information, selects the most favorable trajectory segment for each parameter, and aggregates the resulting scores logarithmically. The optimized trajectory is therefore composed of complementary intervals, each exposing a distinct subset of parameters whose contribution to the motion over that interval can be attributed unambiguously. Across diverse simulation environments, ranging from a linear-dynamics system to the Go2 quadruped, G1 humanoid, and Crazyflie quadrotor, IDTD reduces the parameter identification error by 39.6% on average relative to the strongest prior active-exploration baseline and attains improved downstream policy transfer. Furthermore, we validate the proposed trajectory design on a real K1 humanoid, demonstrating that the resulting identified parameters accurately capture the real-system dynamics.
Commentsarxiv_r1