发表机构
Hybrid Robotics; UC Berkeley(混合机器人实验室; 伯克利大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通过引入轨迹正则化随机最优控制(TRSOC),利用 KL 散度增强标准随机最优控制,推导 HJB 方程刻画最优策略,在 LQ 设置中有闭式解,实验表明正则化参数能在性能与参考保持间权衡。
AI 中文摘要
我们引入了轨迹正则化随机最优控制(TRSOC),它通过受控轨迹分布与参考轨迹分布之间的库尔贝克 - 莱布勒(KL)散度增强了标准随机最优控制(SOC)。利用吉尔萨诺夫定理,轨迹 KL 散度简化为二次漂移失配惩罚,产生了保留动态规划(DP)结构的修正运行成本。我们推导了相应的哈密顿 - 雅可比 - 贝尔曼(HJB)方程并刻画了最优策略。在线性二次(LQ)设置中,该公式允许具有增强控制成本的闭式解。实验表明,正则化参数在性能驱动和参考保持行为之间进行权衡,包括从离线数据学习参考动态的情况。
英文摘要
We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions. Using Girsanov's theorem, the trajectory KL reduces to a quadratic drift mismatch penalty, yielding a modified running cost that preserves the dynamic programming (DP) structure. We derive the corresponding Hamilton--Jacobi--Bellman (HJB) equation and characterize the optimal policy. In the linear-quadratic (LQ) setting, the formulation admits a closed-form solution with an augmented control cost. Experiments show that the regularization parameter induces a trade-off between performance-driven and reference-preserving behavior, including cases with reference dynamics learned from offline data.
Comments8 pages, 4 figures, 65th IEEE Conference on Decision and Control