发表机构
TU Darmstadt; German Research Center for AI (DFKI); Hessian Centre for Artificial Intelligence(达姆施塔特工业大学; 德国人工智能研究中心; 黑森州人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种情境贝叶斯优化框架,利用层级控制器结构联合学习参数,在人形推箱任务中实现最低平均遗憾。
AI 中文摘要
层级控制架构被广泛用于将复杂控制问题分解为相互作用的控制层级,在机器人学中尤为重要,因为规划、全身运动和底层控制必须在不同抽象层次和时间尺度上进行协调。然而,其整体闭环性能在很大程度上依赖于分布在层级中的参数,因此独立调整不同层级的控制器可能会忽略相关的跨层交互。我们提出了一种情境贝叶斯优化框架,用于层级控制系统中的联合参数学习。我们不仅将闭环性能建模为标量黑箱函数,还保留了对任务性能、实现质量和控制努力的单独观测。一个相关的多输出高斯过程对这些性能组件进行建模,而它们已知的聚合到整体闭环目标中的过程则被解析地评估。该公式利用了层级控制的三个互补结果:层级暴露出的信息性性能量、不同控制器层级参数之间的耦合,以及这些关系随操作条件的变化。我们针对人形机器人移动操作评估了该方法,在变化的箱子质量下联合调整一个质心预测控制器和一个全身控制器以进行物理推箱任务。所提出的方法在训练和适应过程中均达到了所考虑基线中最低的平均经验遗憾。
英文摘要
Hierarchical control architectures are widely used to decompose complex control problems into interacting control levels and are particularly important in robotics, where planning, whole-body motion, and lower-level control must be coordinated across different levels of abstraction and time scales. Their overall closed-loop performance, however, depends strongly on parameters distributed across the hierarchy, such that tuning controllers on different levels independently may neglect relevant cross-layer interactions. We propose a contextual Bayesian optimization framework for joint parameter learning in hierarchical control systems. Rather than modeling closed-loop performance only as a scalar black-box function, we retain separate observations of task performance, realization quality, and control effort. A correlated multi-output Gaussian process models these performance components, while their known aggregation into the overall closed-loop objective is evaluated analytically. The formulation exploits three complementary consequences of hierarchical control: informative performance quantities exposed by the hierarchy, coupling between parameters of different controller levels, and variations of these relations with operating conditions. We evaluate the approach for humanoid loco-manipulation, jointly tuning a centroidal predictive controller and a whole-body controller for physical box pushing under varying box mass. The proposed method achieves the lowest mean empirical regret during both training and adaptation among the considered baselines.
Comments8 pages, 4 figures