arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于信号时序逻辑的奖励机

Reward Machines for Signal Temporal Logic

Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai

arXiv 2608.13625首次发表:更新:

发表机构

Virginia Commonwealth University; University of Virginia(弗吉尼亚联邦大学; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对信号时序逻辑控制综合的难题,提出基于定时交替自动机的奖励机方法,在强化学习中实现了更高的鲁棒性分数和规范满足率。

AI 中文摘要

信号时序逻辑(Signal Temporal Logic,STL)提供了一种形式化语言,用于指定实值观测的实时属性,并提供定量鲁棒性分数以监测满足度。由于现实世界系统日益复杂,手动控制器设计变得不可行,因此从STL规范进行控制综合备受关注。此外,许多现代自主和AI驱动的系统缺乏准确完整的系统模型,这使得基于优化的综合方法不适用,从而推动了基于学习的控制的发展。先前的研究将STL鲁棒性分数作为强化学习(Reinforcement Learning,RL)中的奖励,以获得满足给定规范的控制策略;然而,鲁棒性依赖于执行历史,这会导致具有任意嵌套时间算子的一般长程规范的状态空间扩展难以处理。本研究引入了一种新颖的基于自动机的方法,该方法提供了适用于RL框架的高效记忆机制和相关马尔可夫奖励。我们的方法从给定的STL规范构造一个定时交替自动机,通过自动机位置和时钟赋值扩充状态空间,并从自动机接受条件中导出奖励。我们通过实验证明,与现有使用基于鲁棒性的奖励的方法相比,我们的方法学习到的策略实现了更高的鲁棒性分数和满足率。

英文摘要

Signal temporal logic (STL) provides a formal language for specifying real-time properties of real-valued observations, along with a quantitative robustness score for monitoring satisfaction. Control synthesis from STL specifications is of interest since manual controller design becomes infeasible as real-world systems grow in complexity. Moreover, many modern autonomous and AI-enabled systems lack accurate and complete system models, which makes optimization-based synthesis approaches unsuitable and motivates learning-based control. Prior work uses STL robustness scores as rewards in reinforcement learning (RL) to obtain control policies satisfying given specifications; however, robustness depends on execution history, leading to intractable state space expansion for general long-horizon specifications with arbitrarily nested temporal operators. This work introduces a novel automata-based approach that provides an efficient memory mechanism and associated Markovian rewards suitable for RL frameworks. Our approach constructs a timed alternating automaton from the given STL specifications, augments the state space with automaton locations and clock valuations, and derives rewards from the automaton acceptance condition. We empirically demonstrate that our approach learns policies that achieve higher robustness scores and satisfaction rates than those learned by existing approaches using robustness-based rewards.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑