arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用时空管奖励的可处理强化学习:面向全类信号时序逻辑规范

Tractable Reinforcement Learning for Full Class of Signal Temporal Logic Specifications Using Spatiotemporal Tube Reward

Vaishnavi Jagabathula, P Sangeerth, Pushpak Jagtap

arXiv 2609.28396首次发表:更新:

发表机构

Centre for Cyber-Physical Systems, IISc(印度科学研究所网络物理系统中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种时间感知的强化学习框架,利用时空管几何特性将全类信号时序逻辑规范映射为时变边界,训练软演员-评论家智能体,在满足输入约束下实现无历史、计算高效的连续控制策略。

AI 中文摘要

本文针对机器人系统(包括在未知动力学和严格执行器限制下运行的非完整与欠驱动平台)的控制问题,旨在满足复杂的高层规范。我们使用信号时序逻辑(STL)来表示这些高层规范,并提出一种新颖的、时间感知的强化学习(RL)框架,该框架利用时空管(STT)的几何特性。传统的解析STT控制器在强制执行输入约束方面往往存在困难,而现有的RL方法依赖于内存密集型的状态历史,我们的方法则天然地克服了这两个局限性。通过将全类STL的逻辑和时间复杂性映射为时变的几何边界,我们直接约束多维系统状态,而不依赖于标量鲁棒性度量。通过将时间加入状态空间,我们使用一种连续的、几何感知的奖励函数来训练一个时间感知的软演员-评论家(SAC)智能体,该函数无需在执行过程中显式评估复杂的逻辑语义。所提出的框架提供了一种无历史、计算高效的方法,用于学习连续控制策略,确保在严格遵守系统输入约束的同时,稳健地满足规范。

英文摘要

This paper addresses the control problem for robotic systems, including non-holonomic and underactuated platforms operating under unknown dynamics and strict actuator limits to satisfy complex high-level specifications. We denote these high-level specifications using Signal Temporal Logic (STL) and propose a novel time-aware Reinforcement Learning (RL) framework that leverages the geometric properties of Spatiotemporal Tubes (STTs). While traditional analytical STT controllers often struggle to enforce input constraints, and existing RL approaches rely on memory-intensive state history, our method natively overcomes both limitations. By mapping the logical and temporal complexities of the full class of STL into time-varying geometric boundaries, we directly constrain the multidimensional system state without relying on scalar robustness metrics. Augmenting the state space with time, we train a time-aware Soft Actor-Critic (SAC) agent using a continuous, geometry-aware reward function that eliminates the need to explicitly evaluate complex logical semantics during execution. The proposed framework offers a history-free, computationally efficient approach to learn continuous control policies that ensure robust satisfaction of specifications while strictly adhering to system input constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑