哈密尔顿-雅可比可达性中两人博弈价值函数的精确分解
Exact Decomposition of Value Functions for Two-Player Games in Hamilton-Jacobi Reachability
AI总结:
本研究在满足关键单调性假设的前提下,将HJR中适用于单智能体的价值函数分解关键结果推广到两人博弈场景,针对连续时间有限时域任务完成了对应分析。
AI中文摘要:
哈密尔顿-雅可比可达性(HJR)是安全控制理论中的重要框架,它提供了理论工具和数值方法,用于获取涉及目标到达和避障任务的价值函数。近期有研究提出了一种代数框架,通过将复合任务的价值函数分解为HJR中传统研究的基础任务的价值函数,来扩展到复杂的复合任务,且该框架均针对单智能体场景。然而,这些分解结果大多无法直接推广到两人博弈场景。在本技术札记中,我们证明,只要满足关键的单调性假设,上述前期工作中用于分析涉及多个目标到达且需遵守约束的各类任务的关键结果,在两人博弈场景中仍然成立。具体而言,我们详细阐述了对应的结果及其证明(其结构与单智能体场景不同);通过反例展示了若不满足该单调性假设会出现的问题;并说明如何利用该结果分析各类任务。需注意,前期工作考虑的是强化学习中更常见的离散时间、无限时域任务,而本研究考虑的是HJR中更常见的连续时间、有限时域任务。
英文摘要:
Hamilton-Jacobi reachability (HJR) is an important framework in safe control theory. HJR provides theoretical tools and numerical methods to obtain value-functions for tasks involving target-reaching and obstacle-avoidance. A recent work proposed an algebraic framework to scale to complex, composite tasks via decomposing the value functions of the composite tasks into value functions for the fundamental tasks traditionally studied in HJR, all in the one-player setting. Many of these decomposition results, however, do not directly translate to the two-player setting. In this technical note, we nevertheless show that one of the previous work's key results for analyzing various tasks involving reaching multiple targets while obeying constraints still holds in the two-player setting, so long as a critical monotonicity assumption is satisfied. In particular, we detail the analogous result and its proof (which structurally differs from the one-player case), we show via a counter-example what can go wrong if this monotonicity assumption is not satisfied, and we show how the analysis of a variety of tasks can be performed using this result. We note that, whereas the prior work considered discrete-time, infinite-horizon tasks, which are more standard in reinforcement learning, we here consider continuous-time, finite-horizon tasks, which are more standard in HJR.