发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对未知切换动力学的目标跟踪问题,提出自适应性学习结合模型预测控制的方法,经仿真和硬件实验验证其在各类轨迹下的无遗憾跟踪性能优于多种在线学习基线方法。
AI 中文摘要
我们提出了一种用于控制的自适应性在线学习方法,以跟踪未知的目标动力学。目标动力学可表现出切换行为,尤其是结构化、随机和/或对抗性运动的混合。这种具有挑战性的目标跟踪场景出现在动态映射、交通控制和追逃应用中,机器人需要跟踪、追踪或避免与移动地标、物体、人类等发生碰撞,而这些对象的动力学是未知的。我们的方法通过自监督、单次且计算高效的学习,从头开始同时学习多个预测器,并自适应选择最佳预测器以匹配观测到的目标行为。该方法在期望意义上具有有限时间近最优性保证,其特征是作为目标动力学的学习误差和目标动力学切换频率的函数。在没有误差和切换的情况下,该方法渐近匹配最优非因果控制策略(该策略先验已知目标动力学),即该方法在期望意义上具有无遗憾性。在存在学习误差和切换的情况下,该方法会平稳降级,例如当存在误差且无切换时,平均遗憾与平均学习误差和切换次数成正比。为证明这些保证,与采用基于RFF的在线学习的现有工作相比,需要一种新颖的技术方法。我们在Crazyflie仿真和硬件实验中验证了我们的方法,针对从结构化到随机再到对抗性的各种目标轨迹,与非随机、基于核的以及基于神经网络的在线学习方法进行了比较。
英文摘要
We propose a self-adaptive online learning for control method for tracking unknown target dynamics. The target dynamics can exhibit switching behavior, particularly, a mixture of structured, random, and/or adversarial motion. Such challenging target tracking scenarios arise in applications of dynamic mapping, traffic control, and pursuit evasion, where robots need to track, pursue, or avoid collision with moving landmarks, objects, humans, etc., whose dynamics are unknown. Our method simultaneously learns multiple predictors from scratch, via self-supervised, one-shot, and computationally efficient learning, and adaptively selects the best one to match the observed target behavior. The method enjoys finite-time near-optimality guarantees in expectation, characterized as a function of the learning error of the target dynamics and the frequency that the target dynamics switch. In the absence of both error and switching, the method asymptotically matches the optimal non-causal control policy that knows a priori the target dynamics, i.e., the method enjoys no regret in expectation. In the presence of learning errors and switching, the method degrades gracefully, \eg when there are errors and no switching, the average regret is proportional to the average learning error and switching times. To prove these guarantees, a novel technical approach is required compared to the existing works that employ RFF-based online learning. We validate our method in Crazyflie simulations and hardware experiments, across target trajectories that vary from structured to random to adversarial, in comparison to non-stochastic, kernel-based, and neural-network-based methods for online learning.
Comments20 pages, 14 figures