arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.26516cs.LGcs.AI

Lyapunov引导的自对齐:用于离线安全强化学习的测试时适应

Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning

  • Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Seungyub Han, Hyungjin Kim, Jungwoo Lee

更新

AI总结:

本文提出SAS框架,通过自对齐机制在不重新训练的情况下实现离线安全强化学习的测试时适应,利用Lyapunov条件筛选可行轨迹并生成上下文提示,提升安全性与性能。

AI中文摘要:

离线强化学习(RL)代理在部署时往往表现不佳,因为训练数据集与真实环境之间的差距导致不安全行为。为了解决这个问题,我们提出了SAS(用于安全的自对齐),一个基于Transformer的框架,能够在不重新训练的情况下实现离线安全RL的测试时适应。在SAS中,主要机制是自对齐:在测试时,预训练的代理生成多个想象轨迹,并选择满足Lyapunov条件的轨迹。这些可行段落随后被用作上下文提示,使代理能够重新对齐其行为以确保安全,同时避免参数更新。实际上,SAS将Lyapunov引导的想象转化为控制不变的提示,其Transformer架构允许一种层次化RL解释,其中提示功能在潜在技能上进行贝叶斯推断。在Safety Gymnasium和MuJoCo基准测试中,SAS一致地减少了成本和失败,同时保持或提高了回报。

英文摘要:

Offline reinforcement learning (RL) agents often fail when deployed, as the gap between training datasets and real environments leads to unsafe behavior. To address this, we present SAS (Self-Alignment for Safety), a transformer-based framework that enables test-time adaptation in offline safe RL without retraining. In SAS, the main mechanism is self-alignment: at test time, the pretrained agent generates several imagined trajectories and selects those satisfying the Lyapunov condition. These feasible segments are then recycled as in-context prompts, allowing the agent to realign its behavior toward safety while avoiding parameter updates. In effect, SAS turns Lyapunov-guided imagination into control-invariant prompts, and its transformer architecture admits a hierarchical RL interpretation where prompting functions as Bayesian inference over latent skills. Across Safety Gymnasium and MuJoCo benchmarks, SAS consistently reduces cost and failure while maintaining or improving return.

补充信息

↑