发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CF-KKT框架,利用学习动力学和局部最优示范,在不需已知动力学或额外探索的情况下恢复未知约束,结合CIOC与ICRL优势,在高维机器人任务中提升安全性和数据效率。
AI 中文摘要
从示范中学习(LfD)提供了一种从局部最优、满足约束的专家行为中推断未知约束的框架。现有方法主要分为两种范式:约束逆最优控制(CIOC)和逆约束强化学习(ICRL)。CIOC利用最优性条件,如Karush-Kuhn-Tucker(KKT)条件,但通常假设已知动力学和结构化的约束表示。与此同时,ICRL能够处理复杂的未知约束和未知转移动力学,但通常需要大量的在线探索,在此期间可能发生不安全的约束违反。在这项工作中,我们引入了反事实KKT(CF-KKT),一种约束学习框架,利用学习到的动力学和局部最优示范来恢复未知约束,而不需要已知动力学或额外的风险探索,从而结合了CIOC的数据效率和安全性优势以及ICRL的灵活性。首先,我们使用局部学习到的可微动力学模型直接在示范上施加受KKT启发的最优性条件。其次,我们使用学习到的动力学在示范附近生成奖励改进的反事实行为,揭示在不存在未知约束的情况下更可取的行为,从而提供合成的不可行数据。当约束参数化已知时,相同的学习动力学框架能够直接实现基于CIOC的参数恢复,并且我们表征了其对动力学错误指定的敏感性。在高维机器人控制任务中,我们的方法相对于最先进的离线ICRL基线,以更高的安全性和数据效率学习神经约束表示。
英文摘要
Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits optimality conditions such as the Karush--Kuhn--Tucker (KKT) conditions but typically assumes known dynamics and structured constraint representations. Meanwhile, ICRL accommodates complex unknown constraints and unknown transition dynamics but often requires extensive online exploration, during which unsafe constraint violations may occur. In this work, we introduce Counterfactual KKT (CF-KKT), a constraint learning framework that leverages learned dynamics and locally optimal demonstrations to recover unknown constraints without requiring known dynamics or additional risky exploration, thereby combining the data efficiency and safety advantages of CIOC with the flexibility of ICRL. First, we use a locally learned differentiable dynamics model to impose KKT-inspired optimality conditions directly on the demonstrations. Second, we use the learned dynamics to generate reward-improving counterfactual behaviors near the demonstrations, revealing behaviors that would be preferable in the absence of the unknown constraint and thus providing synthetic infeasible data. When the constraint parameterization is known, the same learned-dynamics framework enables direct CIOC-based parameter recovery, and we characterize its sensitivity to dynamics misspecification. Across high-dimensional robotic control tasks, our approach learns neural constraint representations with improved safety and data efficiency relative to state-of-the-art offline ICRL baselines.