Q-learning 惩罚 Transformer 用于安全离线强化学习
Q-learning Penalized Transformer for Safe Offline Reinforcement Learning
- Shanghai Jiao Tong University(上海交通大学)
- Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
- Jilin University(吉林大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 Q-learning 惩罚 Transformer(QPT)框架,通过 Q 形惩罚将安全约束融入条件序列建模,在 DSRL 基准 38 个任务上优于基线并实现零样本阈值适应。
AI中文摘要:
本文研究了安全离线强化学习问题,即利用离线数据集训练策略以满足安全约束。该问题本质上具有挑战性,因为它需要平衡三个高度相互关联且相互竞争的目标:满足安全约束、最大化奖励以及遵循离线数据集施加的行为正则化。为了应对这一三重挑战,我们提出了 Q-learning 惩罚 Transformer 策略(QPT),这是一种“训练-推理一致”的框架,将条件序列建模与约束感知的价值估计相结合。QPT 训练一个 Transformer 策略,该策略根据轨迹上下文和目标回报/代价生成动作,同时保持强大的行为正则化。为了在学习过程中注入明确的安全语义,我们使用学习的奖励和代价 Q 函数,通过 Q 形惩罚增强序列模型训练,以在低约束违反下偏好高回报。在推理时,相同的 Q 函数强制执行代价阈值并选择最高回报的可行动作,从而闭环训练与部署之间的循环。我们在风格化的近确定性 CMDP 下提供了原理性分析,刻画了 Q 惩罚条件生成如何提升安全性和性能。实验上,QPT 在 DSRL 基准的 38 个任务上持续优于强安全离线强化学习基线,并展现出对不同约束阈值的稳健零样本适应能力。
英文摘要:
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizing rewards, and adhering to the behavior regularization imposed by the offline dataset. To tackle this trilogy challenge, we propose Q-learning Penalized Transformer policy (QPT), a \emph{training--inference consistent} framework that bridges conditional sequence modeling with constraint-aware value estimation. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost, retaining strong behavior regularization. To inject explicit safety semantics during learning, we augment sequence-model training with a Q-shaped penalty using learned reward and cost Q-functions to favor high return under low constraint violation. At inference, the same Q-functions enforce the cost threshold and choose the highest-reward feasible action, closing the loop between training and deployment. We provide a principled analysis under stylized near-deterministic CMDPs, characterizing how Q-penalized conditional generation improve safety and performance. Empirically, QPT consistently outperforms strong safe offline RL baselines across 38 tasks on the DSRL benchmark, and exhibits robust zero-shot adaptation to different constraint thresholds.