基于控制障碍函数层的高维安全最优反馈控制的端到端学习
End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers
浏览论文内容
中文总结 AI 辅助
研究在硬安全约束下高维半全局反馈控制器的学习问题,结合算子分裂与无雅可比反向传播方法,实现可扩展端到端训练并保持安全保证,证明该方法在高维多智能体非线性控制问题上的有效性。
中文摘要 AI 辅助
我们考虑在由控制障碍函数(CBF)强制执行的硬安全约束下学习高维半全局反馈控制器的问题。将CBF纳入端到端策略训练需要嵌入基于二次规划的安全滤波器作为优化层,但计算和微分瓶颈很大程度上限制了先前方法应用于低维系统,通常最多16个状态维度。我们通过将算子分裂与最近开发的无雅可比反向传播(JFB)方法相结合来解决这一限制,以实现可扩展的端到端训练,同时通过CBF安全滤波器保持硬安全保证。我们使用非光滑分析技术从理论上证明了这种训练方法的合理性,并在状态和控制维度分别高达1200和400的高维多智能体非线性控制问题上证明了其有效性。
英文摘要
We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier functions (CBFs). Incorporating CBFs into end-to-end policy training requires embedding a quadratic-program-based safety filter as an optimization layer, but computational and differentiation bottlenecks have largely restricted prior approaches to low-dimensional systems, typically with at most 16 state dimensions. We address this limitation by combining operator splitting with the recently developed Jacobian-Free Backpropagation (JFB) method to enable scalable end-to-end training while preserving hard safety guarantees through the CBF safety filter. We justify this training methodology theoretically using nonsmooth analysis techniques and demonstrate its effectiveness on high-dimensional multi-agent nonlinear control problems with state and control dimensions up to 1200 and 400, respectively.